arXiv:2505.07453cs.AI2025-05被引 8

测试大模型处理表格数据的真实推理能力,发现其表现远不如预期。

How well do LLMs reason over tabular data, really?

  • 用LLM当裁判评估模型,比传统方法更可靠
  • 真实表格变化下模型表现大幅下降,缺陷明显
  • 适合关注大模型在数据分析中实际表现的研究者

大型语言模型(LLMs)在自然语言任务中表现出色,但其在表格数据上的推理能力尚不明确。以往的评估策略未能反映模型在真实场景中的表现,且对表格输入中的现实变化(如缺失值、重复实体、结构差异)缺乏充分理解。本文通过一个近期的表格推理基准,揭示了多选题提示评估法及SacreBleu、BERT-score等自由文本指标的不足。研究发现,采用LLM作为评判者的方法能提供更可靠的性能洞察,并暴露了通用大模型在表格推理方面的显著短板。进一步扩展输入数据以包含三种常见实践特征后,实验表明这些变化严重削弱了模型的表现,凸显了提升模型对真实表格输入鲁棒性的必要性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) excel in natural language tasks, but less is known about their reasoning capabilities over tabular data. Prior analyses devise evaluation strategies that poorly reflect an LLM's realistic performance on tabular queries. Moreover, we have a limited understanding of the robustness of LLMs towards realistic variations in tabular inputs. Therefore, we ask: Can general-purpose LLMs reason over tabular data, really?, and focus on two questions 1) are tabular reasoning capabilities of general-purpose LLMs robust to real-world characteristics of tabular inputs, and 2) how can we realistically evaluate an LLM's performance on analytical tabular queries? Building on a recent tabular reasoning benchmark, we first surface shortcomings of its multiple-choice prompt evaluation strategy, as well as commonly used free-form text metrics such as SacreBleu and BERT-score. We show that an LLM-as-a-judge procedure yields more reliable performance insights and unveil a significant deficit in tabular reasoning performance of LLMs. We then extend the tabular inputs reflecting three common characteristics in practice: 1) missing values, 2) duplicate entities, and 3) structural variations. Experiments show that the tabular reasoning capabilities of general-purpose LLMs suffer from these variations, stressing the importance of improving their robustness for realistic tabular inputs.

大模型推理表格数据评估方法鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。