Using AI safely with sensitive language data

What’s the problem? 

Artificial intelligence is rapidly transforming how researchers analyse information, generate code and automate workflows. However, some of the most important language datasets contain highly sensitive information, including transcripts of patient interviews, therapy recordings, learner assessment records, police interviews and legal testimony. 

Privacy legislation, ethics approvals and data governance requirements often prevent this information from being uploaded to cloud-based AI platforms. As a result, researchers working with sensitive language data can be locked out of AI tools that are becoming routine elsewhere. 

 

What’s UQ’s solution? 

The workflow uses a locally hosted AI model running on a researcher’s own computer to generate a synthetic stand-in dataset. This synthetic proxy preserves the structure of the original data but contains entirely fictional content. Researchers can then use cloud-based AI assistants such as ChatGPT or Claude to generate analysis code for that dataset. The resulting code is run locally against the real data, ensuring that sensitive participant information never leaves the researcher’s machine. 

The workflow runs on freely available software, and does not require specialist hardware, making privacy-preserving AI analysis accessible to researchers working with restricted-access datasets. 

What’s the impact? 

The DGUCR workflow enables researchers to benefit from modern AI tools while remaining compliant with privacy legislation, ethics approvals and institutional data governance requirements. 

The approach is designed for researchers and professionals working with restricted-access language data, including working in clinical settings (clinicians, medical researchers, clinical linguists), legal contexts (forensic linguists, legal professionals), education (teacher, second-language researchers), social contexts (social workers, sociolinguists and social scientists). By removing the need to choose between ethics compliance and AI-assisted analysis, the workflow helps make AI safer and more accessible in sensitive research environments. 

Step-by-step tutorials and full code are freely available through LADAL. The workflow has been presented through UQ’s AI in Practice: Lunchtime Insights Series and is being showcased at the 6th Asia Pacific Corpus Linguistics Conference (APCLC2026) in Toyama, Japan. A paper describing the method has been submitted to Research Methods in Applied Linguistics. 

Did you know?

If you follow the DGUCR workflow, the cloud AI never sees the real data. It receives only a synthetic stand-in dataset containing fictional information but the same structure as the original data. Sensitive participant data remains on the researcher’s machine from start to finish. 

Led by

  • Dr Martin Schweinberger 
    Senior Lecturer in Applied Linguistics 
    Director, Language Technology and Data Analysis Laboratory (LADAL) 
    School of Languages and Cultures 

UQ Project Team

Language Technology and Data Analysis Laboratory (LADAL) 

AI Research StrengthHuman-Centred AI 
Industry Portfolio Advancing AI 
Key PartnersLanguage Technology and Data Analysis Laboratory (LADAL), Language Data Commons of Australia (LDaCA) ·  
Key PublicationsTBC
Supporting Linkhttps://ladal.edu.au 

 

99 more AI innovations at UQ