构建真实场景多尺度表格数据集,评估大模型表推理能力
MiMoTable: A Multi-scale Spreadsheet Benchmark with Meta Operations for Table Reasoning
- 基于真实工作表设计多领域表格数据集
- 提出六类元操作评估标准,准确率最高77.4%
- 揭示模型性能随题目难度下降规律,适合评估新基准
现有大语言模型在表格推理任务上表现良好,但与真实应用场景中的复杂表格仍存在显著差距。为此,我们提出一个名为MiMoTable的多尺度电子表格基准,包含七个领域的实际工作表,涵盖多种表格类型。同时,定义了六类元操作作为衡量问题难度的新标准,为现有基准提供新的评估视角。实验表明,Claude-3.5-Sonnet在该基准上取得最高77.4%的准确率,仍有较大提升空间。进一步分析显示,随着基准难度增加,模型性能持续下降,验证了新评估标准的有效性。
原文摘要 · Abstract (English)
Extensive research has been conducted to explore the capability of Large Language Models (LLMs) for table reasoning and has significantly improved the performance on existing benchmarks. However, tables and user questions in real-world applications are more complex and diverse, presenting an unignorable gap compared to the existing benchmarks. To fill the gap, we propose a \textbf{M}ult\textbf{i}-scale spreadsheet benchmark with \textbf{M}eta \textbf{o}perations for \textbf{Table} reasoning, named as MiMoTable. Specifically, MiMoTable incorporates two key features. First, the tables in MiMoTable are all spreadsheets used in real-world scenarios, which cover seven domains and contain different types. Second, we define a new criterion with six categories of meta operations for measuring the difficulty of each question in MiMoTable, simultaneously as a new perspective for measuring the difficulty of the existing benchmarks. Experimental results show that Claude-3.5-Sonnet achieves the best performance with 77.4\% accuracy, indicating that there is still significant room to improve for LLMs on MiMoTable. Furthermore, we grade the difficulty of existing benchmarks according to our new criteria. Experiments have shown that the performance of LLMs decreases as the difficulty of benchmarks increases, thereby proving the effectiveness of our proposed new criterion.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。