用线性探针识别并控制大模型的幻觉生成,效果显著且可解释。
A Single Direction of Truth: An Observer Model's Linear Residual Probe Exposes and Steers Contextual Hallucinations
- 通过残差流中的单一线性方向检测幻觉,无需反向传播。
- 在Gemma-2系列模型上提升5-27分,中层表现稳定。
- 揭示幻觉与特定MLP子电路的因果关联,适合安全可控生成研究者。
上下文幻觉——即与给定上下文无关的陈述——仍是人工智能领域的重大挑战。我们发现一种实用的可解释性方法:仅需一次前向传播和对残差流的线性探针,即可实现生成器无关的幻觉检测。该探针定位到一个单一、可迁移的线性方向,能有效区分幻觉与真实文本,在不同规模的Gemma-2模型(2B至27B)中均表现出色,较基线提升5-27个百分点,并在中层保持稳健性能。梯度乘以激活值分析表明该信号集中于稀疏的晚期MLP活动区域。关键的是,对这一方向的操纵可因果性地调控生成器的幻觉率,证明其可操作性。结果提供了内部低维幻觉追踪机制的新证据,其与特定MLP子电路相关,可用于检测与缓解。我们发布了包含2000个样本的ContraTales基准,用于真实评估此类解决方案。
原文摘要 · Abstract (English)
Contextual hallucinations -- statements unsupported by given context -- remain a significant challenge in AI. We demonstrate a practical interpretability insight: a generator-agnostic observer model detects hallucinations via a single forward pass and a linear probe on its residual stream. This probe isolates a single, transferable linear direction separating hallucinated from faithful text, outperforming baselines by 5-27 points and showing robust mid-layer performance across Gemma-2 models (2B to 27B). Gradient-times-activation localises this signal to sparse, late-layer MLP activity. Critically, manipulating this direction causally steers generator hallucination rates, proving its actionability. Our results offer novel evidence of internal, low-dimensional hallucination tracking linked to specific MLP sub-circuits, exploitable for detection and mitigation. We release the 2000-example ContraTales benchmark for realistic assessment of such solutions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。