What’s the problem?
Many organisations are sitting on a treasure trove of data, but they struggle to turn this into insights. The same information is often stored in tables whose rows and columns sit in different orders, with headers that are inconsistent or missing altogether, which makes it hard to tell when two datasets are describing the same thing. As a result, connections between datasets stay hidden, and duplicated and conflicting records go unnoticed. This can result in poor-quality data that undermine the analyses and AI systems built on top.

What’s UQ’s innovation?
Professor Zhifeng Bao, Ge Lee, and Professor Shazia Sadiq, in UQ’s School of Electrical Engineering and Computer Science, work on turning scattered, messy data into dependable knowledge. They develop methods that automatically identify relevant datasets, uncover the relationships between them, and repair common data-quality problems such as missing values and mismatched records.
Shape-Agnostic Largest Table Overlap (SALTO) is a method that hunts for every truly matching cell between two tables, even when their rows and columns are in a different order and their headers are missing – existing tools could only find the largest neat, rectangular block of matching information two tables share, missing the many matches that fall outside a single tidy block. To make this approach practical at scale, the team developed an algorithm called HyperSplit that finds these overlaps efficiently. The approach could be applied wherever messy tabular data needs to be compared, reconciled or cleaned.
What’s the impact?
The team has shown the approach at work across three data tasks. Comparing daily stock-market tables gathered from 55 different online sources, it pinpointed which sources were copying one another far more accurately than existing tools. This provenance check allows better assessment of data reliability — are multiple sources independently coming to the same conclusion or is a single source feeding many?
Run over Wikipedia’s tables, SALTO located content duplicated across pages through copy-paste and template reuse — the sort of task that helps a large content repository find redundant entries, keep the same fact consistent everywhere it appears, and consolidate to a single source of truth. Comparing successive versions of the same table, it isolated exactly what had changed and what had stayed the same from one revision to the next — useful for any organisation that needs to track how a dataset evolves over time.
More broadly, work like this strengthens the quality and consistency of the data on which downstream analysis and AI systems depend: better inputs make for more trustworthy results. The team’s wider aim is to help people spend less time wrangling data and more time drawing confident conclusions from it.
Did you know?
In tests on real-world datasets, SALTO’s HyperSplit algorithm uncovered more complete matches between tables than the previous best method in up to 78.8 per cent of cases.
Led by
- Professor Zhifeng Bao
ARC Future Fellow, School of Electrical Engineering and Computer Science - Ge Lee
School of Electrical Engineering and Computer Science
PROJECT TEAM
- Professor Shazia Sadiq
School of Electrical Engineering and Computer Science
| AI Research Strength | Data-Centric AI |
|---|---|
| Industry Portfolio | Advancing AI |
| Key Partners | RMIT University University of Wollongong Hasso Plattner Institute (University of Potsdam) CSIRO Data61 |
| Key Publications | Lee, G., Huang, S., Bao, Z., Naumann, F., Sadiq, S. & Zhao, Y. (2026). Shape-Agnostic Table Overlap Discovery: A Maximum Common Subhypergraph Approach. Proceedings of the ACM on Management of Data, 4(3), 1–26. |
| Supporting Link | github.com/DataAutonomyLab/SALTO |
Published 15 September 2026