arXiv:2510.20375cs.CLcs.AI2025-10EMNLP

研究否定语句如何影响大模型幻觉,发现模型在否定文本中更难识别幻觉。

The Impact of Negated Text on Hallucination with Large Language Models

  • 通过重构数据集引入否定表达,测试模型对否定语境的判断能力
  • 模型在否定文本中幻觉检测准确率显著下降,常出现逻辑矛盾
  • 揭示模型处理否定时的内部机制缺陷,适合关注幻觉与推理的研究者

大语言模型(LLMs)的幻觉问题近年来受到广泛关注,但否定语句对幻觉的影响仍缺乏研究。本文提出三个关键未解问题,旨在探究模型是否能识别由否定引发的语境变化,并在否定场景下仍可靠区分幻觉。为此,我们构建了NegHalu数据集,将现有幻觉检测数据集中的陈述改为否定形式。实验表明,LLMs在否定文本中难以有效检测幻觉,常产生逻辑不一致或不忠实的判断。进一步分析模型在词元层面的内部状态,揭示了缓解否定导致的不良影响的内在挑战。

原文摘要 · Abstract (English)

Recent studies on hallucination in large language models (LLMs) have been actively progressing in natural language processing. However, the impact of negated text on hallucination with LLMs remains largely unexplored. In this paper, we set three important yet unanswered research questions and aim to address them. To derive the answers, we investigate whether LLMs can recognize contextual shifts caused by negation and still reliably distinguish hallucinations comparable to affirmative cases. We also design the NegHalu dataset by reconstructing existing hallucination detection datasets with negated expressions. Our experiments demonstrate that LLMs struggle to detect hallucinations in negated text effectively, often producing logically inconsistent or unfaithful judgments. Moreover, we trace the internal state of LLMs as they process negated inputs at the token level and reveal the challenges of mitigating their unintended effects.

幻觉检测大模型否定语义

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。