arXiv:2508.01556cs.AI2025-08ACL被引 5

用大模型提升表格数据准备效率,解决传统方法难题。

Empowering Tabular Data Preparation with Language Models: Why and How?

  • 利用大模型理解表格语义,自动完成数据获取与清洗。
  • 在数据整合与转换阶段显著提升准确率,适配多种下游任务。
  • 适合数据科学家、工程师快速构建高质量数据管道。

数据准备是提升表格数据可用性的关键步骤,直接影响下游数据驱动任务的效果。传统方法难以捕捉表格内部复杂关系且适应性差。近年来,语言模型(特别是大语言模型)的发展为自动化支持表格数据准备提供了新机遇。然而,为何语言模型适用于表格数据准备(即其能力如何匹配任务需求),以及如何在各阶段有效应用仍缺乏系统性探索。本文系统分析了语言模型在增强表格数据准备过程中的作用,聚焦数据获取、集成、清洗和转换四个核心阶段。针对每个阶段,我们整合分析了语言模型与其他组件的结合方式,总结关键进展,并提出未来可行的处理流程。

原文摘要 · Abstract (English)

Data preparation is a critical step in enhancing the usability of tabular data and thus boosts downstream data-driven tasks. Traditional methods often face challenges in capturing the intricate relationships within tables and adapting to the tasks involved. Recent advances in Language Models (LMs), especially in Large Language Models (LLMs), offer new opportunities to automate and support tabular data preparation. However, why LMs suit tabular data preparation (i.e., how their capabilities match task demands) and how to use them effectively across phases still remain to be systematically explored. In this survey, we systematically analyze the role of LMs in enhancing tabular data preparation processes, focusing on four core phases: data acquisition, integration, cleaning, and transformation. For each phase, we present an integrated analysis of how LMs can be combined with other components for different preparation tasks, highlight key advancements, and outline prospective pipelines.

大模型数据准备表格数据LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。