arXiv:2510.21118cs.CLcs.AI2025-10被引 2

提出新标注框架,解决大模型摘要幻觉检测中的模糊性问题

The Gray Zone of Faithfulness: Taming Ambiguity in Unfaithfulness Detection

  • 引入中间类别'Out-Dependent'区分需外部知识验证的摘要
  • 发现SOTA模型如GPT-5仍有约6%句子存在幻觉
  • 适用于评估大模型摘要忠实性与改进幻觉检测方法

确保大语言模型(LLMs)生成的摘要忠实于源文档对实际应用至关重要。现有基准因生成结果中允许使用外部知识的边界不明确而存在标注模糊问题。例如常识常被纳入回应并标记为'忠实',但其合理范围未明确定义,导致标注不一致。为此,我们提出一种新的忠实性标注框架,引入中间类别'Out-Dependent',用于分类需要外部知识才能验证的情况。基于此框架,我们构建了VeriGray(带灰区验证的未忠实性检测基准)——一项新的摘要任务未忠实性检测基准。统计显示,即使是最先进的模型如GPT-5,在摘要任务中仍存在约6%的幻觉句子。此外,平均约9%的生成句子属于'Out-Dependent'类别,凸显解决标注模糊性的必要性。实验表明,该基准对多种基线方法构成显著挑战,揭示了未来改进空间。

原文摘要 · Abstract (English)

Ensuring that Large Language Models (LLMs) generate summaries faithful to a given source document is essential for real-world applications. While prior research has explored LLM faithfulness, existing benchmarks suffer from annotation ambiguity, primarily due to the ill-defined boundary of permissible external knowledge in generated outputs. For instance, common sense is often incorporated into responses and labeled as "faithful", yet the acceptable extent of such knowledge remains unspecified, leading to inconsistent annotations. To address this issue, we propose a novel faithfulness annotation framework, which introduces an intermediate category, Out-Dependent, to classify cases where external knowledge is required for verification. Using this framework, we construct VeriGray (Verification with the Gray Zone) -- a new unfaithfulness detection benchmark in summarization. Statistics reveal that even SOTA LLMs, such as GPT-5, exhibit hallucinations ($\sim 6\%$ of sentences) in summarization tasks. Moreover, a substantial proportion ($\sim 9\%$ on average of models) of generated sentences fall into the Out-Dependent category, underscoring the importance of resolving annotation ambiguity in unfaithfulness detection benchmarks. Experiments demonstrate that our benchmark poses significant challenges to multiple baseline methods, indicating considerable room for future improvement.

大模型摘要忠实性幻觉检测标注框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。