通过放大模型内部特征,让语言模型自己解释隐藏表示。
Superscopes: Amplifying Internal Feature Representations for Language Model Interpretation
- 利用特征方向放大机制增强弱但有意义的内部特征。
- 无需额外训练即可揭示此前方法无法解释的复杂概念表示。
- 适合研究大模型工作机制与可解释性的研究人员。
理解大型语言模型(LLMs)的内部表征仍是一个开放性挑战。Patchscopes 提出通过将内部激活值替换到新提示中,使模型自我解释其隐藏表征。我们提出 Superscopes,一种系统性放大多层感知机(MLP)输出和隐藏状态中叠加特征的技术,在进行上下文替换前增强这些特征。受“特征即方向”视角及扩散模型中的无分类器引导(CFG)方法启发,Superscopes 能够放大微弱但具有意义的特征,从而实现对以往方法无法解释的内部表征的解读,且无需额外训练。该方法为理解语言模型如何构建上下文和表示复杂概念提供了新视角,进一步推动了机制可解释性的发展。
原文摘要 · Abstract (English)
Understanding and interpreting the internal representations of large language models (LLMs) remains an open challenge. Patchscopes introduced a method for probing internal activations by patching them into new prompts, prompting models to self-explain their hidden representations. We introduce Superscopes, a technique that systematically amplifies superposed features in MLP outputs (multilayer perceptron) and hidden states before patching them into new contexts. Inspired by the "features as directions" perspective and the Classifier-Free Guidance (CFG) approach from diffusion models, Superscopes amplifies weak but meaningful features, enabling the interpretation of internal representations that previous methods failed to explain-all without requiring additional training. This approach provides new insights into how LLMs build context and represent complex concepts, further advancing mechanistic interpretability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。