在大模型中用稀疏自编码器提取可解释特征,发现包含抽象概念和潜在危害行为的神经模式。
Scaling Monosemanticity: Extracting Interpretable Features from Claude 3 Sonnet

- 用稀疏自编码器从大模型中间层提取3400万特征,基于缩放定律优化超参数
- 提取特征涵盖名人、地点、讽刺、代码错误及欺骗等抽象与有害概念
- 这些特征可直接影响模型输出,适合用于可解释性研究与安全对齐
我们证明稀疏自编码器可从生产级语言模型Claude 3 Sonnet中提取可解释特征,解决了字典学习方法能否扩展到大模型的开放问题。在模型中间层残差流上训练了最多达3400万特征的稀疏自编码器,并利用缩放定律指导超参数选择。所得特征具备多语言、多模态特性(尽管仅在文本上训练,仍能泛化至图像),对具体实例和抽象概念均有响应,并可用于按语义引导模型行为。识别出著名人物、地点,以及讽刺、代码错误等抽象概念的特征,还发现了代表欺骗、权力追求、阿谀奉承和偏见等可能造成危害的特征,并证实其在操纵时会因果影响模型输出。此外,我们进行了特征可解释性、几何结构与计算功能分析。但仍有显著局限:特征集不完整,且缺乏严格评估方法以验证特征是否忠实反映模型计算。
原文摘要 · Abstract (English)
We demonstrate that sparse autoencoders can extract interpretable features from Claude 3 Sonnet, a production-scale language model, addressing the open question of whether dictionary learning methods scale beyond small transformers. We trained sparse autoencoders with up to 34 million features on the model's middle layer residual stream, using scaling laws to guide hyperparameter selection. The resulting features are multilingual and multimodal (generalizing to images despite text-only training), respond to both concrete instances and abstract discussions of concepts, and can be used to steer model behavior in ways consistent with their interpretations. We find features corresponding to famous entities and locations, as well as more abstract concepts like sarcasm or errors in code. We also identify features relevant to ways in which language models might cause harm--including features representing deception, power-seeking, sycophancy, and bias--and show that these causally influence model outputs when manipulated. Additionally, we conduct analyses of feature interpretability, geometry, and computational function. However, significant limitations remain: our suite of features is incomplete, and we lack rigorous methods for evaluating whether our features faithfully capture model computations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。