arXiv:2505.22176cs.CL2025-05ACL被引 4

提出可解释的表格评估框架,精准识别表格细微错误。

TabXEval: Why this is a Bad Table? An eXhaustive Rubric for Table Evaluation

  • 分两阶段评估:先结构对齐,再语义语法对比
  • 在多领域真实扰动数据集上验证,表现稳定可靠
  • 适合需要高精度表格生成与质检的研究者

表格的定性与定量评估面临重大挑战,因标准指标常忽略细微的结构和内容差异。为此,我们提出基于评分表的评估框架,融合多层次结构描述符与细粒度上下文信号,实现更精确、一致的表格比较。在此基础上,我们构建了TabXEval——一个全面且可解释的两阶段评估框架。首先通过TabAlign对齐参考表与预测表的结构,再利用TabCompare进行语义与语法层面的比较,提供可解释的细粒度反馈。我们在包含真实表格扰动与人工标注的多领域基准集TabXBench上评估了TabXEval。敏感性-特异性分析进一步验证了其在各类表格任务中的鲁棒性与可解释性。代码与数据已公开于 https://coral-lab-asu.github.io/tabxeval/

原文摘要 · Abstract (English)

Evaluating tables qualitatively and quantitatively poses a significant challenge, as standard metrics often overlook subtle structural and content-level discrepancies. To address this, we propose a rubric-based evaluation framework that integrates multi-level structural descriptors with fine-grained contextual signals, enabling more precise and consistent table comparison. Building on this, we introduce TabXEval, an eXhaustive and eXplainable two-phase evaluation framework. TabXEval first aligns reference and predicted tables structurally via TabAlign, then performs semantic and syntactic comparison using TabCompare, offering interpretable and granular feedback. We evaluate TabXEval on TabXBench, a diverse, multi-domain benchmark featuring realistic table perturbations and human annotations. A sensitivity-specificity analysis further demonstrates the robustness and explainability of TabXEval across varied table tasks. Code and data are available at https://coral-lab-asu.github.io/tabxeval/

表格评估可解释性基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。