arXiv:2506.08249cs.DBcs.CL2025-06被引 5

评测大模型在含缺陷表格数据上的分析能力,发现其表现大幅下降。

RADAR: Benchmarking Language Models on Imperfect Tabular Data

  • 通过程序化扰动生成9类真实数据中的5种缺陷,构建可控制的评测框架
  • 2980组表格查询对显示,有缺陷数据下前沿模型性能显著下滑
  • 适合关注数据质量、模型鲁棒性与真实场景应用的研究者使用

语言模型(LMs)正被越来越多地用于自主数据分析,但其对数据缺陷(如缺失值、异常值、逻辑不一致)的认知与处理能力仍缺乏系统研究。这些缺陷在真实世界表格数据中极为常见,若处理不当将严重影响分析结论的有效性。为此,我们提出RADAR,一个系统评估模型在表格数据上进行数据感知推理能力的基准。通过程序化扰动构建框架,模拟9个领域中5种类型的数据缺陷,涵盖2980组表格-查询对。除评估缺陷处理能力外,还系统变化表格规模以考察推理性能随数据量增长的变化。实验表明,尽管在无缺陷表格上表现良好,前沿模型在引入数据缺陷后性能急剧下降,暴露出其在稳健、数据感知分析方面的关键短板。RADAR设计灵活可扩展,支持多种扰动类型和可控表规模,为推进表格推理研究提供重要资源。

原文摘要 · Abstract (English)

Language models (LMs) are increasingly being deployed to perform autonomous data analyses. However, their data awareness -- the ability to recognize, reason over, and appropriately handle data artifacts such as missing values, outliers, and logical inconsistencies -- remains underexplored. These artifacts are especially common in real-world tabular data and, if mishandled, can significantly compromise the validity of analytical conclusions. To address this gap, we present RADAR, a benchmark for systematically evaluating data-aware reasoning on tabular data. We develop a framework to simulate data artifacts via programmatic perturbations to enable targeted evaluation of model behavior. RADAR comprises 2980 table query pairs, grounded in real-world data spanning 9 domains and 5 data artifact types. In addition to evaluating artifact handling, RADAR systematically varies table size to study how reasoning performance holds when increasing table size. Our evaluation reveals that, despite decent performance on tables without data artifacts, frontier models degrade significantly when data artifacts are introduced, exposing critical gaps in their capacity for robust, data-aware analysis. Designed to be flexible and extensible, RADAR supports diverse perturbation types and controllable table sizes, offering a valuable resource for advancing tabular reasoning.

大模型数据质量评测基准表格推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。