构建表格推理与鲁棒性评测基准,揭示当前模型在表格任务上表现脆弱。
The Mighty ToRR: A Benchmark for Table Reasoning and Robustness
- 构建涵盖10个数据集的表格推理与鲁棒性评测基准ToRR
- 强模型在不同表格格式下表现不稳,存在显著脆弱性
- 多格式测试和多提示验证能显著提升评估可靠性
尽管表格数据具有重要现实意义,但模型在表格任务上的表现仍缺乏系统评估,导致难以判断应采用何种模型或提示配置。为此,我们提出了ToRR基准,用于衡量模型在表格相关任务中的性能与鲁棒性。该基准包含10个数据集,覆盖多种领域和表格推理能力。ToRR不仅关注模型性能排名,更强调模型在不同常见表格表示格式下的稳定性与一致性。我们发布了排行榜,并对主流模型在ToRR上的结果进行了全面分析。结果显示,模型行为极为脆弱,即使强模型也难以在表格任务中保持稳健。尽管特定表格格式未带来持续优势,但多格式测试对于可靠评估模型能力至关重要。此外,多提示验证带来的可靠性提升相当于增加更多测试样本。总体表明,表格理解与推理仍是重大挑战。
原文摘要 · Abstract (English)
Despite its real-world significance, model performance on tabular data remains underexplored, leaving uncertainty about which model to rely on and which prompt configuration to adopt. To address this gap, we create ToRR, a benchmark for Table Reasoning and Robustness, measuring model performance and robustness on table-related tasks. The benchmark includes 10 datasets that cover different types of table reasoning capabilities across varied domains. ToRR goes beyond model performance rankings, and is designed to reflect whether models can handle tabular data consistently and robustly, across a variety of common table representation formats. We present a leaderboard as well as comprehensive analyses of the results of leading models over ToRR. Our results reveal a striking pattern of brittle model behavior, where even strong models are unable to perform robustly on tabular data tasks. Although no specific table format leads to consistently better performance, we show that testing over multiple formats is crucial for reliably estimating model capabilities. Moreover, we show that the reliability boost from testing multiple prompts can be equivalent to adding more test examples. Overall, our findings show that table understanding and reasoning tasks remain a significant challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。