arXiv:2504.10284cs.CL2025-04ACL综述被引 8

构建真实评估基准,提升大模型生成文献综述表格的可靠性。

arXiv2Table: Toward Realistic Benchmarking and Evaluation for LLM-Based Literature-Review Table Generation

  • 设计无金标准泄漏的用户需求与干扰论文,模拟真实检索噪声。
  • 提出分解式评估指标,涵盖列覆盖、单元准确性与关系一致性。
  • 适用于需要可复现评估的AI辅助文献综述研究者。

文献综述表格对归纳和对比科学论文至关重要。本文研究从论文集合中自动生成满足用户信息需求的表格。基于近期工作(Newman等,2024),我们突破理想化设定:(i) 模拟明确但无模式依赖的用户需求,避免泄露黄金列名或值;(ii) 通过人工验证的语义相关但不相关论文引入检索噪声;(iii) 提出轻量级、无需标注、以利用率为导向的评估方法,将实用性分解为模式覆盖率、单个单元准确性和成对关系一致性,并通过双向问答流程(金标到系统、系统到金标)测量论文选择,使用召回率、精确率和F1值。为支持可复现评估,我们构建了arXiv2Table基准,包含1,957张表格,参考7,158篇论文,含人工验证的干扰项和重写的无模式用户需求。我们还开发了迭代式批量生成方法,多轮协同优化论文筛选与模式定义。通过人工审计和交叉评估验证评估协议。大量实验表明,该方法持续优于强基线,但绝对得分仍较低,凸显任务难度。数据与代码见https://github.com/JHU-CLSP/arXiv2Table。

原文摘要 · Abstract (English)

Literature review tables are essential for summarizing and comparing collections of scientific papers. In this paper, we study the automatic generation of such tables from a pool of papers to satisfy a user's information need. Building on recent work (Newman et al., 2024), we move beyond oracle settings by (i) simulating well-specified yet schema-agnostic user demands that avoid leaking gold column names or values, (ii) explicitly modeling retrieval noise via semantically related but out-of-scope distractor papers verified by human annotators, and (iii) introducing a lightweight, annotation-free, utilization-oriented evaluation that decomposes utility into schema coverage, unary cell fidelity, and pairwise relational consistency, while measuring paper selection through a two-way QA procedure (gold to system and system to gold) with recall, precision, and F1. To support reproducible evaluation, we introduce arXiv2Table, a benchmark of 1,957 tables referencing 7,158 papers, with human-verified distractors and rewritten, schema-agnostic user demands. We also develop an iterative, batch-based generation method that co-refines paper filtering and schema over multiple rounds. We validate the evaluation protocol with human audits and cross-evaluator checks. Extensive experiments show that our method consistently improves over strong baselines, while absolute scores remain modest, underscoring the task's difficulty. Our data and code is available at https://github.com/JHU-CLSP/arXiv2Table.

文献综述表格生成评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。