arXiv:2509.00869cs.CL2025-09被引 8

提出对比解码方法,有效减少大模型因迎合误导输入而产生的虚假回应。

Exploring and Mitigating Fawning Hallucinations in Large Language Models

  • 设计两种诱导误导输入的范式,精准触发模型的奉承幻觉。
  • 无需额外训练,通过对比输出分布显著提升生成内容的真实性。
  • 适用于多种自然语言任务,适合关注模型可信度的研究者使用。

大型语言模型在语言理解方面表现卓越,但当其输出与欺骗性或误导性提示对齐时,生成内容可能偏离事实。这种现象称为奉承幻觉,即模型优先迎合输入隐含立场而非真实信息。本文分析了不同自然语言处理任务中的奉承幻觉,并提出协同对比解码(CCD)方法以缓解该问题。具体地,设计两种范式生成对应的欺骗性/误导性输入,以诱发一致的奉承幻觉;随后,通过对比诱导输入与转换后的中性输入在输出分布上的差异,实现不依赖额外训练的幻觉抑制。大量实验表明,所提方法能有效减轻奉承幻觉,提升生成内容的事实性,在多个任务上均取得显著效果。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated exceptional proficiency in language understanding. However, when LLMs align their outputs with deceptive and/or misleading prompts, the generated responses could deviate from the de facto information. Such observations are known as fawning hallucinations, where the model prioritizes alignment with the input's implied perspective over accuracy and truthfulness. In this work, we analyze fawning hallucinations in various natural language processing tasks and tailor the so-termed contrastive decoding method for fawning-hallucination mitigation. Specifically, we design two paradigms to generate corresponding deceptive and/or misleading inputs for the consistent fawning hallucinations induction. Then, we propose the collaborative contrastive decoding (CCD) to handle the fawning hallucinations across different tasks in LLMs. By contrasting the deviation in output distribution between induced and transformed neutral inputs, the proposed CCD can reduce reliance on deceptive and/or misleading information without requiring additional training. Extensive experiments demonstrate that the proposed CCD can effectively mitigate fawning hallucinations and improve the factuality of the generated responses over various tasks.

大模型幻觉对比解码事实性增强

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。