arXiv:2505.15291cs.CL2025-05被引 10

发现大模型生成长文本时后半部分易幻觉,提出针对性缓解策略。

Hallucinate at the Last in Long Response Generation: A Case Study on Long Document Summarization

  • 分析长文档摘要中幻觉在输出序列中的位置分布规律。
  • 发现幻觉集中出现在生成文本后半部分,比例超预期。
  • 针对结尾幻觉设计优化方法,提升长文本结尾忠实度。

大型语言模型在文本生成任务中表现出色,但在忠实于源材料方面仍面临幻觉问题。尽管已有大量研究致力于检测与减少错误信息,但对幻觉在长输出中位置分布的关注较少。本文以长文档摘要为关键案例,研究了长上下文感知的长文本生成中幻觉的位置特征。结果表明,在长响应生成中,幻觉往往集中在生成内容的后半部分,呈现显著偏倚。我们进一步探究了注意力机制和解码过程在长序列中的动态特性对这一现象的影响,并提出针对性缓解方法,旨在提升长文本生成结尾部分的忠实性。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have significantly advanced text generation capabilities, including tasks like summarization, often producing coherent and fluent outputs. However, faithfulness to source material remains a significant challenge due to the generation of hallucinations. While extensive research focuses on detecting and reducing these inaccuracies, less attention has been paid to the positional distribution of hallucination within generated text, particularly in long outputs. In this work, we investigate where hallucinations occur in LLM-based long response generation, using long document summarization as a key case study. Focusing on the challenging setting of long context-aware long response generation, we find a consistent and concerning phenomenon: hallucinations tend to concentrate disproportionately in the latter parts of the generated long response. To understand this bias, we explore potential contributing factors related to the dynamics of attention and decoding over long sequences. Furthermore, we investigate methods to mitigate this positional hallucination, aiming to improve faithfulness specifically in the concluding segments of long outputs.

幻觉检测长文本生成摘要

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。