The year/Independent research

Paper 2601.17058

Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs

Published
Jan 2026
Research lab
Independent
Citations
6
GitHub
814 stars

01 In brief

Summary

This paper surveys the use of large language models (LLMs) for data preparation, covering data cleaning, integration, and enrichment.

It contrasts traditional rule-based and model-specific methods with LLM-enhanced approaches that leverage prompting, retrieval-augmented generation (RAG), fine-tuning, and agentic workflows.

The survey identifies three core tasks: data cleaning (standardization, error processing, imputation), data integration (entity matching, schema matching), and data enrichment (annotation, profiling).

For each, it reviews representative techniques, highlighting strengths such as improved generalization and semantic understanding, and limitations like high inference costs, hallucinations, and evaluation mismatches.

It also summarizes common datasets and evaluation metrics, and discusses open challenges and future directions, including scalable LLM-data systems, reliable agentic workflows, and robust evaluation protocols.

The paper emphasizes a paradigm shift from manual, rule-based pipelines to prompt-driven, context-aware, and agentic preparation workflows, and notes trends toward hybrid LLM-ML methods, reduced fine-tuning, and cross-modal generalization.

02 From the paper

Abstract

Data preparation aims to denoise raw datasets, uncover cross-dataset relationships, and extract valuable insights from them, which is essential for a wide range of data-centric applications. Driven by (i) rising demands for application-ready data (e.g., for analytics, visualization, decision-making), (ii) increasingly powerful LLM techniques, and (iii) the emergence of infrastructures that facilitate flexible agent construction (e.g., using Databricks Unity Catalog), LLM-enhanced methods are rapidly becoming a transformative and potentially dominant paradigm for data preparation. By investigating hundreds of recent literature works, this paper presents a systematic review of this evolving landscape, focusing on the use of LLM techniques to prepare data for diverse downstream tasks. First, we characterize the fundamental paradigm shift, from rule-based, model-specific pipelines to prompt-driven, context-aware, and agentic preparation workflows. Next, we introduce a task-centric taxonomy that organizes the field into three major tasks: data cleaning (e.g., standardization, error processing, imputation), data integration (e.g., entity matching, schema matching), and data enrichment (e.g., data annotation, profiling). For each task, we survey representative techniques, and highlight their respective strengths (e.g., improved generalization, semantic understanding) and limitations (e.g., the prohibitive cost of scaling LLMs, persistent hallucinations even in advanced agents, the mismatch between advanced methods and weak evaluation). Moreover, we analyze commonly used datasets and evaluation metrics (the empirical part). Finally, we discuss open research challenges and outline a forward-looking roadmap that emphasizes scalable LLM-data systems, principled designs for reliable agentic workflows, and robust evaluation protocols.