arXiv:2608.10172cs.LG2026-08

提出可识别的模型内在谱特征,解决解释性方法结果不可靠问题。

Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability

论文配图:Intrinsic Structure: Spectral Identifiability for Mechanistic Interpretability
图 1 · 摘自论文原文
  • 用柯普曼算子将模型前向过程视为动力系统,提取不变谱特征。
  • 在多个大模型上验证谱收敛,误差率符合理论预测(0.506±0.031)。
  • 揭示可识别特征与可读分解本质不同,适合追求理论可信度的研究者。

机制可解释性通过识别模型内部回路来解释模型行为,但无法判断回路是模型固有属性还是方法产物。稀疏自编码器即体现此问题:不同随机种子和宽度会从相同激活中恢复出显著不同的特征,且无理论说明这种差异是偶然还是结构性的。本文将解释性字典学习置于可识别性框架下:将前向传播视为以深度为时间的受控动力系统,通过柯普曼算子提升后得到有限线性实现,其谱为模型的坐标无关属性。我们证明该谱可在M个校准样本下以速率M^{-1/2}恢复(至排列意义),据我们所知为首个机制可解释性基本构件的可识别性定理,包含匹配极小极大下界、重尾激活的中位数-均值变体及解耦定理:当实现非正规时,承载激活方差的方向与跨深度传递信息的方向不可能重合。可识别对象与可读对象并非同一实体。在GPT-2 small、Gemma-2-2B和Qwen3-8B-Base上,谱全局收敛,且在Qwen3-8B-Base上达到预测指数0.506±0.031;偏差集中于各单元样本阈值的统一曲线上。柯普曼模式优于随机方向但弱于主成分,在间接目标识别中差距随深度距离衰减4.1倍,符合理论预期。柯普曼谱是带有明确误差条的模型内在指纹,而非可读分解。

原文摘要 · Abstract (English)

Mechanistic interpretability explains models by identifying circuits inside them, but has no way to tell whether a circuit is a property of the model or an artifact of the method that found it. Sparse autoencoders illustrate the problem: different seeds and widths recover materially different features from the same activations, and no theory says whether that variability is incidental or structural. We put dictionary learning for interpretability on an identifiability footing. Treating the forward pass as a controlled dynamical system with depth as time and lifting it with the Koopman operator yields a finite linear realisation whose \emph{spectrum} is a coordinate-free property of the model. We prove the spectrum is recoverable from $M$ calibration samples at rate $M^{-1/2}$ up to permutation - to our knowledge the first identifiability theorem for a mechanistic-interpretability primitive, with a matching minimax lower bound, a median-of-means variant for heavy-tailed activations, and a dissociation theorem: whenever the realisation is non-normal, the directions carrying activation variance and the directions carrying information across depth cannot coincide. The identifiable object and the legible object are not the same object. On GPT-2 small, Gemma-2-2B and Qwen3-8B-Base the spectrum converges everywhere and attains the predicted exponent on Qwen3-8B-Base ($0.506 \pm 0.031$); shortfalls collapse onto one curve against each cell's sample threshold. Koopman modes beat random directions but lose to principal components on indirect-object identification, with the gap decaying $4.1\times$ in depth-distance, as the theorem predicts. The Koopman spectrum is an identifiable, model-intrinsic fingerprint with a stated error bar, not a legible decomposition.

可解释性谱分析模型指纹机制解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。