arXiv:2507.16488cs.CLcs.AI2025-07ACL被引 36

通过追踪隐藏状态动态变化,精准识别大模型幻觉

ICR Probe: Tracking Hidden State Dynamics for Reliable Hallucination Detection in LLMs

  • 提出ICR分数衡量模块对隐藏状态更新的贡献
  • 在多个数据集上检测准确率超基线15%以上
  • 参数量少、可解释性强,适合部署于资源受限场景

大型语言模型在自然语言处理任务中表现优异,但其生成幻觉问题严重影响可靠性。现有基于隐藏状态的检测方法多关注静态孤立表征,忽略其跨层动态演变,导致效果受限。本文聚焦隐藏状态更新过程,提出全新指标ICR Score(信息对残差流的贡献),量化模块对隐藏状态更新的贡献。实验证明ICR Score能有效区分幻觉与真实内容。基于此,我们设计ICR Probe方法,捕捉隐藏状态的跨层演化特性。结果表明,该方法在保持高精度的同时,参数量显著减少。消融实验与案例分析揭示了其内在机制,提升了可解释性。

原文摘要 · Abstract (English)

Large language models (LLMs) excel at various natural language processing tasks, but their tendency to generate hallucinations undermines their reliability. Existing hallucination detection methods leveraging hidden states predominantly focus on static and isolated representations, overlooking their dynamic evolution across layers, which limits efficacy. To address this limitation, we shift the focus to the hidden state update process and introduce a novel metric, the ICR Score (Information Contribution to Residual Stream), which quantifies the contribution of modules to the hidden states' update. We empirically validate that the ICR Score is effective and reliable in distinguishing hallucinations. Building on these insights, we propose a hallucination detection method, the ICR Probe, which captures the cross-layer evolution of hidden states. Experimental results show that the ICR Probe achieves superior performance with significantly fewer parameters. Furthermore, ablation studies and case analyses offer deeper insights into the underlying mechanism of this method, improving its interpretability.

幻觉检测隐藏状态LLM可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。