arXiv:2605.00939cs.LGcs.AI2026-05中稿 · ICML

通过梯度敏感性检测大模型高自信下的顽固错误。

From Flat Facts to Sharp Hallucinations: Detecting Stubborn Errors via Gradient Sensitivity

论文配图:From Flat Facts to Sharp Hallucinations: Detecting Stubborn Errors via Gradient Sensitivity
图 1 · 摘自论文原文
  • 用噪声扰动输入嵌入,观察梯度突增来判断答案是否脆弱。
  • 在多个数据集上比熵值和表征方法更准确识别错误。
  • 适合需要可靠事实验证的场景,如医疗、法律领域应用。

传统幻觉检测在应对‘顽固幻觉’时失效——即大模型自信地犯错。本文提出几何解法:嵌入扰动梯度敏感性(EPGS)。假设稳定事实存在于平坦极小值区域,而顽固幻觉则位于尖锐极小值区,依赖脆弱记忆。EPGS通过向输入嵌入添加高斯噪声,测量梯度幅值的突增,以此作为海塞谱的有效代理,区分稳定知识与不稳定记忆。实验表明,EPGS显著优于基于熵和表征的基线方法,在多个数据集上提供了鲁棒的高置信度事实错误检测信号。

原文摘要 · Abstract (English)

Traditional hallucination detection fails on "Stubborn Hallucinations" - errors where LLMs are confidently wrong. We propose a geometric solution: Embedding-Perturbed Gradient Sensitivity (EPGS). We hypothesize that while robust facts reside in flat minima, stubborn hallucinations sit in sharp minima, supported by brittle memorization. EPGS detects this sharpness by perturbing input embeddings with Gaussian noise and measuring the resulting spike in gradient magnitude. This acts as an efficient proxy for the Hessian spectrum, differentiating stable knowledge from unstable memorization. Our experiments show that EPGS significantly outperforms entropy-based and representation-based baselines, providing a robust signal for detecting high-confidence factual errors.

幻觉检测梯度分析大模型可靠性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。