大模型让数据清洗更智能,自动处理脏数据、整合多源信息并丰富数据内涵。
Can LLMs Clean Up Your Mess? A Survey of Application-Ready Data Preparation with LLMs
- 用提示词驱动和智能体方式替代传统规则流程,实现上下文感知的数据准备。
- 在数据清洗、集成与丰富任务中展现更强语义理解能力,但存在幻觉与成本高问题。
- 适合关注数据自动化处理、想提升数据质量的研究者与从业者参考。
数据准备旨在去噪原始数据集、挖掘跨数据集关联并提取有价值洞见,对数据分析、可视化与决策等应用至关重要。随着对应用就绪数据需求上升、大语言模型技术日益强大,以及支持灵活智能体构建的基础设施(如 Databricks Unity Catalog)出现,基于大模型的数据准备正快速成为主流范式。本文系统综述了近年数百篇相关研究,聚焦大模型在多样化下游任务中的数据准备应用。首先,揭示从规则驱动、模型特定流程向提示驱动、上下文感知、智能体化工作流的根本转变;其次,提出以任务为中心的分类体系,涵盖数据清洗(如标准化、错误处理、缺失值填补)、数据集成(如实体匹配、模式匹配)和数据增强(如标注、画像)三大类。针对每类任务,总结代表性方法,分析其优势(如泛化性提升、语义理解增强)与局限(如大模型扩展成本高、先进智能体仍存幻觉、方法与评估手段不匹配)。同时梳理常用数据集与评估指标。最后,探讨开放挑战,提出面向可扩展的 LLM-数据系统、可靠智能体工作流设计及稳健评估协议的发展路线图。
原文摘要 · Abstract (English)
Data preparation aims to denoise raw datasets, uncover cross-dataset relationships, and extract valuable insights from them, which is essential for a wide range of data-centric applications. Driven by (i) rising demands for application-ready data (e.g., for analytics, visualization, decision-making), (ii) increasingly powerful LLM techniques, and (iii) the emergence of infrastructures that facilitate flexible agent construction (e.g., using Databricks Unity Catalog), LLM-enhanced methods are rapidly becoming a transformative and potentially dominant paradigm for data preparation. By investigating hundreds of recent literature works, this paper presents a systematic review of this evolving landscape, focusing on the use of LLM techniques to prepare data for diverse downstream tasks. First, we characterize the fundamental paradigm shift, from rule-based, model-specific pipelines to prompt-driven, context-aware, and agentic preparation workflows. Next, we introduce a task-centric taxonomy that organizes the field into three major tasks: data cleaning (e.g., standardization, error processing, imputation), data integration (e.g., entity matching, schema matching), and data enrichment (e.g., data annotation, profiling). For each task, we survey representative techniques, and highlight their respective strengths (e.g., improved generalization, semantic understanding) and limitations (e.g., the prohibitive cost of scaling LLMs, persistent hallucinations even in advanced agents, the mismatch between advanced methods and weak evaluation). Moreover, we analyze commonly used datasets and evaluation metrics (the empirical part). Finally, we discuss open research challenges and outline a forward-looking roadmap that emphasizes scalable LLM-data systems, principled designs for reliable agentic workflows, and robust evaluation protocols.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。