arXiv:2506.18421cs.CLcs.AI2025-06被引 3

构建首个全面的表格推理评测基准,揭示大模型在真实场景下的短板

TReB: A Comprehensive Benchmark for Evaluating Table Reasoning Capabilities of Large Language Models

  • 按浅层理解与深层推理分层设计26项子任务
  • 基于高质量数据集,三类推理模式验证模型能力
  • 开源数据集与评测框架,助力模型优化

企业与行业中的大部分数据以表格、数据库和数据仓库形式存储。由于表格数据具有隐含语义、内在复杂性和结构化特征,大语言模型(LLMs)在处理表格推理任务时面临巨大挑战。其中一个重要问题在于缺乏一个能公平反映模型广泛表格推理能力的有效评估基准。本文通过提出TReB——一个全面的表格推理评测基准,填补了这一空白。首先,我们构建了一个分类体系,系统性地衡量浅层表格理解与深层表格推理能力,涵盖26项子任务;随后,通过专门的数据处理与合成流程构建高质量数据集;基于这些精心构造的样本,设计了三种不同推理模式的评估框架,以稳健测量表格推理能力。实验结果表明,现有大模型在应对复杂且真实的表格相关任务时仍存在显著提升空间。数据集与评估框架均已公开,数据集托管于https://huggingface.co/datasets/JT-LM/JIUTIAN-TReB,框架代码见https://github.com/JT-LM/jiutian-treb。

原文摘要 · Abstract (English)

The majority of data in businesses and industries is stored in tables, databases, and data warehouses. Reasoning with table-structured data poses significant challenges for large language models (LLMs) due to its hidden semantics, inherent complexity, and structured nature. One of these challenges is lacking an effective evaluation benchmark fairly reflecting the performances of LLMs on broad table reasoning abilities. In this paper, we fill in this gap by presenting a comprehensive table reasoning benchmark, TReB. Firstly, we propose a taxonomy to systematically measure both shallow table understanding abilities and deep table reasoning abilities, covering a total of 26 sub-tasks. We then construct a high quality dataset through a dedicated data processing and synthesis procedure. Based on these well-constructed samples, we design an evaluation framework to robustly measure table reasoning capabilities with three distinct inference modes. Experimental results with our data and framework reveal that existing LLMs still have significant room for improvement in addressing the complex and real world table related tasks. Both the dataset and evaluation framework are publicly available, with the dataset hosted on https://huggingface.co/datasets/JT-LM/JIUTIAN-TReB, and the framework on https://github.com/JT-LM/jiutian-treb.

表格推理评测基准大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。