arXiv:2607.06482cs.CLcs.AI2026-07

评测大模型在真实复杂数据中的分析能力,发现现有模型差距显著。

Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities

论文配图:Data Analysis in the Wild: Benchmarking Large Language Models Against Real-World Data Complexities
图 1 · 摘自论文原文
  • 基于政府公开数据构建真实场景评测基准
  • 模型在复杂问题与洞察发现任务中表现不佳
  • 适合研究数据智能与大模型应用的学者参考

当前评估大语言模型(LLM)数据分析能力的基准往往无法反映真实场景。它们通常聚焦于小表格的事实检索,忽视了大型多表数据集、外部知识整合以及探索性洞察发现等挑战。我们提出DataGovBench,一个源自政府开放数据的基准,用于评估模型在实际应用场景中的表现。该基准包含两项任务:Table QA要求解决复杂可分解问题并生成文本答案或可视化结果;Table Insight评估模型通过探索性数据分析生成专家级发现的能力。对前沿大模型(含代理框架)的全面实验揭示了两项任务中均存在显著性能差距。结果表明,当前基于大模型的系统仍远未满足真实数据智能需求。DataGovBench为提升模型分析与洞察能力的研究提供了挑战性基准。代码与示例数据见https://github.com/SoHasegawa/datagovbench。

原文摘要 · Abstract (English)

Current benchmarks for evaluating Large Language Models (LLMs) in data analysis often fail to reflect real-world settings. They typically focus on fact retrieval from small tables and overlook the challenges of large multi-tabular datasets, external knowledge integration, and exploratory insight discovery. We introduce DataGovBench, a benchmark derived from governmental open data designed to evaluate LLMs in practical scenarios. The benchmark includes two tasks: Table QA that requires solving complex decomposable questions and producing textual answers or visualizations, and Table Insight that evaluates the ability of models to generate expert-level findings through exploratory data analysis. Comprehensive experiments with state-of-the-art LLMs, both with and without agentic frameworks, reveal significant performance gaps across both tasks. These results suggest that current LLM-based systems remain far from satisfying the demands of real-world data analytics. DataGovBench provides a challenging benchmark for advancing research on LLMs capable of both answering analytical queries and discovering insights from data. Code and sample data are available at https://github.com/SoHasegawa/datagovbench.

大模型评测数据洞察真实数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。