arXiv:2605.21849cs.LGcs.CL2026-05

提出可自适应调整的解释器,提升模型在分布外情况下的解释可靠性。

Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift

论文配图:Geometry-Adaptive Explainer for Faithful Dictionary-Based Interpretability under Distribution Shift
图 1 · 摘自论文原文
  • 基于几何自适应机制,动态对齐分布外激活子空间与解释词典。
  • 在多个模型和分布外场景下,解释忠实度超越所有训练型基线。
  • 仅需无标签分布外数据,无需反向传播,适合实际部署场景。

机制可解释性旨在通过识别模型内部因果结构来解释其行为。基于词典的解释方法(如稀疏自编码器和转换器)是主要工具,但其在分布外(OOD)情形下的忠实度尚未得到系统研究。我们发现,分布偏移会旋转模型实际使用的子空间,导致在分布内(ID)激活上训练的解释词典发生错位。我们将这种错位形式化为‘忠实度差距’,即ID词典与分布外活跃子空间之间的几何距离,并证明该距离控制了分布外忠实度的下降。为此,我们提出几何自适应解释器(GAE),在不改变原始特征结构的前提下,仅使用无标签分布外激活数据,将解释词典重新对齐至分布外活跃子空间。我们证明,相较于未适配的ID解释器,GAE的额外损失以二阶矩偏移为界,呈二次增长。实验表明,在多个模型与分布外设置中,GAE的因果忠实度甚至超过所有基于训练的基线。

原文摘要 · Abstract (English)

Mechanistic interpretability aims to explain a model's behavior by identifying causally responsible internal structures. Dictionary-based explainers such as sparse autoencoders and transcoders are a primary tool, but their faithfulness under out-of-distribution (OOD) shift has received little systematic attention. We show that distribution shift rotates the subspace that the model actively uses, misaligning the explainer's dictionary trained on in-distribution (ID) activations. We formalize this misalignment as the faithfulness gap, a geometric distance between the ID dictionary and the OOD-active subspace, and show that it controls OOD faithfulness degradation. To reduce this gap, we propose the Geometry-Adaptive Explainer (GAE), which realigns the explainer's dictionary with the OOD-active subspace while preserving the original feature structure. This requires only unlabeled OOD activations and no gradient updates. We prove that GAE improves over the unadapted ID explainer, with excess loss bounded quadratically by the second-moment shift. Empirically, GAE even matches or surpasses all training-based baselines in causal faithfulness across multiple models and OOD settings.

可解释性分布外词典解释几何自适应

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。