提出评估长文档摘要中论点覆盖度的新框架,解决法律与科研文本摘要遗漏关键论据的问题。
ARC: Argument Representation and Coverage Analysis for Zero-Shot Long Document Summarization with Instruction Following LLMs
- 从论点类型和信息缺失类型两方面区分评估摘要质量
- 发现模型在论点稀疏分布时易遗漏关键信息,且存在位置偏见
- 适用于法律、科研等高风险领域摘要的改进,适合关注生成可靠性研究者
我们提出论点表征覆盖率(ARC),一种自下而上的评估框架,用于衡量摘要对关键论点的保留程度,这一问题在法律等高风险领域尤为重要。ARC通过区分需覆盖的信息类型,并将遗漏与事实错误分离,提供可解释的评估视角。利用该框架,我们在论点角色核心的两个领域——长篇法律意见书和科学论文中,评估了八种开源大模型的摘要表现。结果表明,尽管模型能捕捉部分重要论点角色,但在论点分布稀疏时频繁遗漏关键信息。此外,ARC揭示出系统性模式:上下文窗口的位置偏倚和角色特定偏好显著影响论点覆盖率,并为构建更完整可靠的摘要策略提供了可操作指导。
原文摘要 · Abstract (English)
We introduce Argument Representation Coverage (ARC), a bottom-up evaluation framework that assesses how well summaries preserve salient arguments, a crucial issue in summarizing high-stakes domains such as law. ARC provides an interpretable lens by distinguishing between different information types to be covered and by separating omissions from factual errors. Using ARC, we evaluate summaries from eight open-weight large language models in two domains where argument roles are central: long legal opinions and scientific articles. Our results show that while these models capture some salient roles, they frequently omit critical information, particularly when arguments are sparsely distributed across the input. Moreover, ARC uncovers systematic patterns, showing how context window positional bias and role-specific preferences shape argument coverage, and provides actionable guidance for developing more complete and reliable summarization strategies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。