arXiv:2604.16430cs.CLcs.AI2026-04被引 2

用稀疏自编码器捕捉大模型幻觉的动态特征,定位错误源头。

HalluSAE: Detecting Hallucinations in Large Language Models via Sparse Auto-Encoders

  • 基于能量景观建模生成过程,识别幻觉发生的关键过渡区。
  • 在Gemma-2-9B上实现当前最佳幻觉检测效果,准确率显著提升。
  • 适合关注模型可解释性与事实性验证的研究者使用。

大型语言模型虽强大且广泛应用,但其实际效果受限于著名的幻觉现象。尽管近期检测方法已有进展,我们发现多数方法忽视了幻觉的动态特性及其内在机制。为此,我们提出HalluSAE,一个受相变启发的框架,将幻觉视为模型潜在动态中的关键转变。通过将生成过程建模为势能景观中的轨迹,HalluSAE识别出临界转变区域,并将事实错误归因于特定高能稀疏特征。该方法包含三个阶段:(1) 利用稀疏自编码器和几何势能度量定位相位区域;(2) 使用对比逻辑归因法识别与幻觉相关的稀疏特征;(3) 通过对解耦特征进行线性探测,实现因果幻觉检测。在Gemma-2-9B上的大量实验表明,HalluSAE达到当前最优的幻觉检测性能。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are powerful and widely adopted, but their practical impact is limited by the well-known hallucination phenomenon. While recent hallucination detection methods have made notable progress, we find most of them overlook the dynamic nature and underlying mechanisms of it. To address this gap, we propose HalluSAE, a phase transition-inspired framework that models hallucination as a critical shift in the model's latent dynamics. By modeling the generation process as a trajectory through a potential energy landscape, HalluSAE identifies critical transition zones and attributes factual errors to specific high-energy sparse features. Our approach consists of three stages: (1) Potential Energy Empowered Phase Zone Localization via sparse autoencoders and a geometric potential energy metric; (2) Hallucination-related Sparse Feature Attribution using contrastive logit attribution; and (3) Probing-based Causal Hallucination Detection through linear probes on disentangled features. Extensive experiments on Gemma-2-9B demonstrate that HalluSAE achieves state-of-the-art hallucination detection performance.

幻觉检测稀疏自编码器可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。