arXiv:2606.07617cs.LGcs.AI2026-06中稿 · ICML

Query Lens通过联合分析键值特征,揭示稀疏特征的激活与影响。

Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects

论文配图:Query Lens: Interpreting Sparse Key-Value Features with Indirect Effects
图 1 · 摘自论文原文
  • 联合编码器键特征与解码器值特征,全面识别特征激活源与输出影响。
  • 发现Logit Lens无法解释的特征在Query Lens下呈现一致的词元签名。
  • 提出子空间通道假说,解释下游模块如何通过特定子空间读取特征。

稀疏自编码器提供的特征虽比单个神经元更易解释,但可靠表征仍具挑战。本文提出Query Lens,将Logit Lens扩展为更全面、更忠实的稀疏特征解释工具。通过联合分析编码器侧的键特征与解码器侧的值特征,识别出激活某一特征的输入及其促进的输出。同时,考虑特征经下游模块处理后产生的间接、模块中介效应,超越了Logit Lens仅捕捉直接效应的局限。实验表明,Query Lens使原本在Logit Lens下无法解释的特征呈现出一致的词元签名。最后,我们提出子空间通道假说,认为下游模块通过层特定子空间读取特征。

原文摘要 · Abstract (English)

While sparse autoencoders provide features more interpretable than individual neurons, reliably characterizing them remains challenging. We propose Query Lens, which extends Logit Lens to enable more comprehensive and faithful interpretations of sparse features. By jointly considering encoder-side key features and decoder-side value features, we identify both the inputs that activate a feature and the outputs it promotes. We also account for indirect, module-mediated effects that arise when the feature is processed by downstream modules, going beyond the direct effect captured by Logit Lens. In experiments, we find that Query Lens yields coherent token signatures for features that remain uninterpretable under Logit Lens. Finally, we propose the Subspace Channel Hypothesis, suggesting that downstream modules read features through layer-specific subspaces.

模型解释稀疏编码特征分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。