arXiv:2606.26620cs.LGcs.AI2026-06

用稀疏自编码器挖掘指令微调模型的可解释特征

Discovering Millions of Interpretable Features with Sparse Autoencoders

论文配图:Discovering Millions of Interpretable Features with Sparse Autoencoders
图 1 · 摘自论文原文
  • 在多个Qwen3模型层上训练稀疏自编码器,提取可解释特征
  • 发现不同层与组件存在独特的稀疏性-保真度权衡
  • 特征可因果调控模型拒绝行为,适合机制研究者使用

稀疏自编码器(SAEs)已成为分解语言模型表征为稀疏且可解释特征的强大工具。然而,训练SAEs计算成本高,现有开源模型仍有限。本文提出 extbf{Qwen3-Instruct SAE},一套基于Qwen3指令微调模型族(Qwen3-1.7B、Qwen3-4B、Qwen3-8B)训练的SAE系统。对Qwen3-1.7B和Qwen3-4B,在残差流、MLP输出和注意力输出三个关键激活位置进行逐层训练;对Qwen3-8B,在部分残差流层上训练。通过激活级重建与模型级恢复指标系统评估,揭示了各层与组件间的稀疏性-保真度权衡。最后,通过拒绝引导案例研究,证明所选SAE特征可因果调控指令微调的Qwen3模型产生拒绝行为。本发布为研究稀疏表征、特征级机制与行为干预提供了实用资源。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have emerged as a powerful tool for decomposing superposed language model representations into sparse and interpretable features. However, training SAEs is computationally expensive, and available open-source SAE models remain limited. In this work, we introduce \textbf{Qwen3-Instruct SAE}, a comprehensive suite of SAEs trained on the Qwen3 instruction-tuned model family, covering Qwen3-1.7B, Qwen3-4B, and Qwen3-8B. For Qwen3-1.7B and Qwen3-4B, we train layer-wise SAEs at three key activation sites: residual streams, MLP outputs, and attention outputs. For Qwen3-8B, we train SAEs on a subset of residual stream layers. We systematically evaluate these SAEs using both activation-level reconstruction metrics and model-level recovery metrics, revealing distinct sparsity--fidelity trade-offs across layers and components. Finally, we demonstrate the utility of Qwen3-Instruct SAE through a refusal-steering case study, showing that selected SAE features can causally steer instruction-tuned Qwen3 models toward refusal behavior. Our release provides a practical resource for studying sparse representations, feature-level mechanisms, and behavioral interventions in instruction-tuned language models

稀疏编码可解释性语言模型特征工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。