arXiv:2412.13378cs.CL2024-12被引 1

用可执行编辑评估摘要中的事实错误检测与解释能力

SummExecEdit: A Factual Consistency Benchmark in Summarization with Executable Edits

  • 通过可执行编辑构建事实一致性评测框架
  • 顶尖模型联合得分仅0.49,检测与解释分别0.67和0.73
  • 超半数大模型在30%以上任务中表现不佳,多类解释错误

事实一致性检测在摘要生成中至关重要,但现有基准缺乏足够挑战性和可解释性。本文提出SummExecEdit,一个结合可执行编辑的新型评测流水线与基准,用于评估模型在发现事实错误并提供准确解释方面的能力。在该基准上,表现最佳的模型Claude3-Opus的联合检测与解释得分为0.49,其中检测得分0.67,解释得分0.73。我们对20多个大语言模型进行了详细评估,发现超过一半模型在30%以上的测试样本中表现不佳。此外,我们识别出四类主要解释错误,其中45.4%的错误表现为关注与摘要内容完全无关的部分。

原文摘要 · Abstract (English)

Detecting factual inconsistencies in summarization is critical, yet existing benchmarks lack the necessary challenge and interpretability for robust evaluation. In this paper, we introduce SummExecEdit, a novel pipeline and benchmark leveraging executable edits to assess models on their ability to both detect factual errors and provide accurate explanations. The top-performing model, Claude3-Opus, achieves a joint detection and explanation score of only 0.49 in our benchmark, with individual scores of 0.67 for detection and 0.73 for explanation. We conduct detailed evaluations to assess the current state of models in this field and find that more than half of the 20+ LLMs in our study struggle with over 30% of the SummExecEdit benchmark. Additionally, we identify four primary types of explanation errors, with 45.4% of them involving a focus on completely unrelated parts of the summary.

摘要评估事实一致性大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。