Finding the hidden connections in messy data

What’s the problem? 

Many organisations are sitting on a treasure trove of data, but they struggle to turn this into insights. The same information is often stored in tables whose rows and columns sit in different orders, with headers that are inconsistent or missing altogether, which makes it hard to tell when two datasets are describing the same thing.  As a result, connections between datasets stay hidden, and duplicated and conflicting records go unnoticed. This can result in poor-quality data that undermine the analyses and AI systems built on top. 

SALTO was used to compare stock-market tables from different online sources to determine which ones were copying one another. Photo credit: Adobe Stock

 

What’s UQ’s innovation? 

Professor Zhifeng Bao, Ge Lee, and Professor Shazia Sadiq, in UQ’s School of Electrical Engineering and Computer Science, work on turning scattered, messy data into dependable knowledge. They develop methods that automatically identify relevant datasets, uncover the relationships between them, and repair common data-quality problems such as missing values and mismatched records. 

Shape-Agnostic Largest Table Overlap (SALTO) is a method that hunts for every truly matching cell between two tables, even when their rows and columns are in a different order and their headers are missing – existing tools could only find the largest neat, rectangular block of matching information two tables share, missing the many matches that fall outside a single tidy block. To make this approach practical at scale, the team developed an algorithm called HyperSplit that finds these overlaps efficiently. The approach could be applied wherever messy tabular data needs to be compared, reconciled or cleaned. 

What’s the impact? 

The team has shown the approach at work across three data tasks. Comparing daily stock-market tables gathered from 55 different online sources, it pinpointed which sources were copying one another far more accurately than existing tools. This provenance check allows better assessment of data reliability — are multiple sources independently coming to the same conclusion or is a single source feeding many? 

Run over Wikipedia’s tables, SALTO located content duplicated across pages through copy-paste and template reuse — the sort of task that helps a large content repository find redundant entries, keep the same fact consistent everywhere it appears, and consolidate to a single source of truth. Comparing successive versions of the same table, it isolated exactly what had changed and what had stayed the same from one revision to the next — useful for any organisation that needs to track how a dataset evolves over time. 

More broadly, work like this strengthens the quality and consistency of the data on which downstream analysis and AI systems depend: better inputs make for more trustworthy results. The team’s wider aim is to help people spend less time wrangling data and more time drawing confident conclusions from it. 

Did you know?

In tests on real-world datasets, SALTO’s HyperSplit algorithm uncovered more complete matches between tables than the previous best method in up to 78.8 per cent of cases. 

Led by

  • Professor Zhifeng Bao  
    ARC Future Fellow, School of Electrical Engineering and Computer Science 
  • Ge Lee
    School of Electrical Engineering and Computer Science
     

PROJECT TEAM 

AI Research StrengthData-Centric AI  
Industry Portfolio Advancing AI
Key Partners

RMIT University 

University of Wollongong 

Hasso Plattner Institute (University of Potsdam) 

CSIRO Data61 

Key PublicationsLee, G., Huang, S., Bao, Z., Naumann, F., Sadiq, S. & Zhao, Y. (2026). Shape-Agnostic Table Overlap Discovery: A Maximum Common Subhypergraph Approach. Proceedings of the ACM on Management of Data, 4(3), 1–26.  
Supporting Linkgithub.com/DataAutonomyLab/SALTO 

Published 15 September 2026

99 more AI innovations at UQ