让大模型真正用好真实世界的表格数据,解决噪声和复杂结构问题
Toward Real-World Table Agents: Capabilities, Workflows, and Design Principles for LLM-based Table Intelligence
- 定义五项核心能力,系统评估大模型处理表格的能力
- 发现开源模型在真实场景中表现显著落后于学术测试
- 提出实用设计原则,提升模型在真实业务中的鲁棒性
表格在金融、医疗和公共管理等领域至关重要,但现实任务常面临噪声、结构异质性和语义复杂性等问题,现有研究多聚焦于干净的学术数据集,未充分覆盖真实场景。本文聚焦基于大模型的表格智能代理(LLM-based Table Agents),旨在通过整合预处理、推理与领域适配,自动化表格为中心的工作流程。提出五项核心能力:C1表结构理解、C2表与查询语义理解、C3表检索与压缩、C4可执行推理与可追溯性、C5跨域泛化,用于分析和比较现有方法。对Text-to-SQL代理的深入分析显示,学术基准与真实场景间存在显著性能差距,尤其在开源模型上。最后,提供切实可行的建议,以提升表格智能代理在实际应用中的鲁棒性、泛化能力和效率。
原文摘要 · Abstract (English)
Tables are fundamental in domains such as finance, healthcare, and public administration, yet real-world table tasks often involve noise, structural heterogeneity, and semantic complexity--issues underexplored in existing research that primarily targets clean academic datasets. This survey focuses on LLM-based Table Agents, which aim to automate table-centric workflows by integrating preprocessing, reasoning, and domain adaptation. We define five core competencies--C1: Table Structure Understanding, C2: Table and Query Semantic Understanding, C3: Table Retrieval and Compression, C4: Executable Reasoning with Traceability, and C5: Cross-Domain Generalization--to analyze and compare current approaches. In addition, a detailed examination of the Text-to-SQL Agent reveals a performance gap between academic benchmarks and real-world scenarios, especially for open-source models. Finally, we provide actionable insights to improve the robustness, generalization, and efficiency of LLM-based Table Agents in practical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。