arXiv:2511.07318cs.CLcs.AI2025-11被引 3

模型误把数据偏见当真相,导致幻觉检测失效。

When Bias Pretends to Be Truth: How Spurious Correlations Undermine Hallucination Detection in LLMs

  • 利用训练数据中的虚假关联(如姓氏与国籍)生成错误但自信的回答。
  • 现有检测方法在虚假关联下完全失效,且不随模型增大而改善。
  • 适合关注大模型可靠性、幻觉机制与公平性研究的读者。

尽管取得显著进展,大语言模型仍会生成看似合理实则错误的回应(即幻觉)。本文揭示了一类此前未被充分关注的幻觉:由数据中的虚假相关性驱动——训练数据中特征(如姓氏)与属性(如国籍)之间存在表面显著但无实际因果关系的统计关联。我们通过系统性合成实验和对主流开源及专有模型(包括GPT-5)的实证评估发现,这类虚假相关性会导致模型生成高度自信的幻觉,不受模型规模影响,逃避当前检测手段,且在拒绝微调后依然持续存在。理论分析进一步说明,这些统计偏见从根本上破坏了基于置信度的检测机制。研究强调必须开发专门应对虚假相关性引发幻觉的新方法。

原文摘要 · Abstract (English)

Despite substantial advances, large language models (LLMs) continue to exhibit hallucinations, generating plausible yet incorrect responses. In this paper, we highlight a critical yet previously underexplored class of hallucinations driven by spurious correlations -- superficial but statistically prominent associations between features (e.g., surnames) and attributes (e.g., nationality) present in the training data. We demonstrate that these spurious correlations induce hallucinations that are confidently generated, immune to model scaling, evade current detection methods, and persist even after refusal fine-tuning. Through systematically controlled synthetic experiments and empirical evaluations on state-of-the-art open-source and proprietary LLMs (including GPT-5), we show that existing hallucination detection methods, such as confidence-based filtering and inner-state probing, fundamentally fail in the presence of spurious correlations. Our theoretical analysis further elucidates why these statistical biases intrinsically undermine confidence-based detection techniques. Our findings thus emphasize the urgent need for new approaches explicitly designed to address hallucinations caused by spurious correlations.

幻觉检测虚假相关大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。