用新方法分解小模型参数,找出可解释的语义单元
Decomposition of Small Transformer Models
- 改进因果重要性函数,适配序列数据的参数分解
- 在GPT-2-small中定位到'高尔夫''篮球'等可解释组件
- 为理解现代模型内部机制提供可干预的分析路径
近期机制可解释性研究显示,对参数空间进行分解能为分析与干预提供清晰接口。以往方法虽在各类简化模型中成功应用,但尚未跨越到真实模型。本文将随机参数分解(SPD)扩展至Transformer模型,提出适用于序列数据的新因果重要性函数与损失函数。实验表明,SPD可成功分解一个模拟归纳头模型,并恢复预期的两步电路;同时在GPT-2-small中成功定位对应于'高尔夫'、'篮球'等可解释概念的子组件。这些结果标志着将SPD拓展至现代模型的第一步,证明该方法可用于揭示参数空间中的可解释机制。
原文摘要 · Abstract (English)
Recent work in mechanistic interpretability has shown that decomposing models in parameter space may yield clean handles for analysis and intervention. Previous methods have demonstrated successful applications on a wide range of toy models, but the gap to "real models" has not yet been bridged. In this work, we extend Stochastic Parameter Decomposition (SPD) to Transformer models, proposing an updated causal importance function suited for sequential data and a new loss function. We demonstrate that SPD can successfully decompose a toy induction-head model and recover the expected 2-step circuit. We also show that applying SPD to GPT-2-small can successfully locate subcomponents corresponding to interpretable concepts like "golf" and "basketball". These results take the first step in the direction of extending SPD to modern models, and show that we can use the method to surface interpretable parameter-space mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。