构建可扩展的表格数学推理数据集,揭示模型在真实表格上的脆弱性
TabularMath: Understanding Math Reasoning over Tables with Large Language Models
- 用自动转换框架生成可控的表格推理任务
- 发现表格复杂度与质量共同影响模型表现,低质表危害严重
- 适合研究大模型表格推理能力或鲁棒性的研究人员
数学推理长期是评估大语言模型的关键基准。尽管数学应用题已有显著进展,但现实应用中对表格数据的多步数值推理需求被忽视。例如,商业智能不仅需要处理表格中的多步计算,还需应对信息不全或不一致的问题。然而,该领域评估严重受限,主要依赖难以扩展的人工收集表格,且缺乏对真实场景中潜在陷阱的覆盖。为此,我们提出AutoT2T神经符号框架,可可控地将数学应用题转化为可扩展且可验证的表格推理任务。基于此流程,我们构建了TabularMath基准,包含四个子集,涵盖文本和图像表格,覆盖表格复杂度、质量及表示形式维度。研究发现:(1)表格复杂度与推理难度共同影响性能;(2)低质量表格对当前大模型推理构成严重风险;(3)不同表格模态趋势相似,文本表格通常更易处理。针对每项发现进行深入分析,为未来研究提供指引。
原文摘要 · Abstract (English)
Mathematical reasoning has long been a key benchmark for evaluating large language models. Although substantial progress has been made on math word problems, the need for reasoning over tabular data in real-world applications has been overlooked. For instance, applications such as business intelligence demand not only multi-step numerical reasoning with tables but also robustness to incomplete or inconsistent information. However, comprehensive evaluation in this area is severely limited, constrained by the reliance on manually collected tables that are difficult to scale and the lack of coverage for potential traps encountered in real-world scenarios. To address this problem, we propose AutoT2T, a neuro-symbolic framework that controllably transforms math word problems into scalable and verified tabular reasoning tasks. Building on this pipeline, we develop TabularMath, a benchmark comprising four subsets that include both text-based and image-based tables, covering table complexity, table quality, and table representation dimensions. Our study reveals three key observations: (1) Table complexity and reasoning difficulty impact reasoning performance jointly; (2) Low-quality tables pose severe risks to reliable reasoning in current LLMs; (3) Different table modalities show similar trends, with text-based tables typically being easier for models to reason over. In-depth analyses are conducted for each observation to guide future research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。