arXiv:2511.05745cs.LGcs.AI2025-11被引 1

提升大模型可解释性的同时降低计算成本,让专家网络各司其职。

Beyond Redundancy: Diverse and Specialized Multi-Expert Sparse Autoencoder

  • 多专家并行激活+特征自适应缩放,促使专家学习不同功能。
  • 重建误差降低24%,特征冗余减少99%。
  • 适合需要高效且精准解释大模型的开发者与研究人员。

稀疏自编码器(SAEs)已成为解析大语言模型(LLMs)的重要工具,通过将标记激活分解为人类可理解的特征组合实现模型可解释性。然而,高维隐藏层以满足稀疏性约束导致训练与推理成本过高,制约了实际应用。现有混合专家(MoE)方法通过分组窄专家网络并采用门控激活来缓解计算压力,但存在关键缺陷:专家缺乏专业化,常学习重叠或相同特征。为此,本文提出两项创新:(1) 多专家激活机制,同时启用语义加权的专家子集以促进分工;(2) 特征缩放机制,通过自适应高频缩放增强特征多样性。实验表明,相比现有MoE-SAE方法,本方法重建误差降低24%,特征冗余减少99%。该工作弥合了大模型分析中可解释性与效率之间的鸿沟,实现在计算可行前提下的透明模型审查。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) have emerged as a powerful tool for interpreting large language models (LLMs) by decomposing token activations into combinations of human-understandable features. While SAEs provide crucial insights into LLM explanations, their practical adoption faces a fundamental challenge: better interpretability demands that SAEs' hidden layers have high dimensionality to satisfy sparsity constraints, resulting in prohibitive training and inference costs. Recent Mixture of Experts (MoE) approaches attempt to address this by partitioning SAEs into narrower expert networks with gated activation, thereby reducing computation. In a well-designed MoE, each expert should focus on learning a distinct set of features. However, we identify a \textit{critical limitation} in MoE-SAE: Experts often fail to specialize, which means they frequently learn overlapping or identical features. To deal with it, we propose two key innovations: (1) Multiple Expert Activation that simultaneously engages semantically weighted expert subsets to encourage specialization, and (2) Feature Scaling that enhances diversity through adaptive high-frequency scaling. Experiments demonstrate a 24\% lower reconstruction error and a 99\% reduction in feature redundancy compared to existing MoE-SAE methods. This work bridges the interpretability-efficiency gap in LLM analysis, allowing transparent model inspection without compromising computational feasibility.

可解释性稀疏编码混合专家大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。