arXiv:2608.00205cs.CL2026-08

人类标注的摘要忠实性常因局部错误被误判为可信,存在平均偏差。

Averaging Bias: Human Faithfulness Annotations are not Locally Faithful

论文配图:Averaging Bias: Human Faithfulness Annotations are not Locally Faithful
图 1 · 摘自论文原文
  • 用大模型逐句评分,发现人类全局标注更倾向平均表现而非全句正确。
  • 约30%被标记为忠实的摘要含真实事实错误,说明标注标准宽松。
  • 适用于评估人类标注可靠性及改进评测设计的研究者参考。

文本摘要忠实性评估采用严格合取规则:只有当摘要每句话都有源文档支持时才视为忠实。然而,现有基准大多仅收集每条摘要一个全局人类标注。本文研究发现,人类标注与大模型逐句评分的平均值相关性更高,而非严格满足合取规则。手动审查显示,约30%被标为忠实的摘要包含真实局部事实错误。这种倾向称为「平均偏差」。结果表明,当前广泛使用的忠实性基准中的人类标注存在可测量的平均偏差,亟需更严谨的标注设计以确保评估可信度。

原文摘要 · Abstract (English)

Evaluation of faithfulness of text summarization treats a model generated summary as faithful only if every of its sentences is supported by the source document: a strict conjunctive rule under which a single unsupported sentence makes the whole summary unfaithful. Yet most faithfulness benchmarks collect only one global human annotation label per summary. We ask whether such global human labels actually implement the conjunctive rule. We hypothesize that annotators may accept a summary as faithful when most sentences are faithful, not only when all are faithful. To test our hypothesis, we use five large language model (LLM) judges as per-sentence raters across four widely used faithfulness benchmarks. We find that global human labels correlate better with the average of per-sentence LLM judgments than with the implementation of the strict conjunctive rule. A manual review confirms that a substantial fraction of summaries labeled faithful by humans contain genuine local factual errors. We call this tendency Averaging Bias. Our results reveal that human labels on widely used faithfulness benchmarks contain measurable Averaging Bias, calling for carefully structured designs for trustworthy human annotations

忠实性评估人类标注平均偏差评测设计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。