arXiv:2512.05534cs.LGcs.AI2025-12被引 2

首次建立稀疏字典学习统一理论,解释模型为何出现特征混淆与失效。

A Unified Theory of Sparse Dictionary Learning in Mechanistic Interpretability: Piecewise Biconvexity and Spurious Minima

  • 将多种稀疏字典学习方法统一为分段双凸优化问题
  • 揭示特征吸收与神经元失效的理论根源,精度提升超40%
  • 适合研究可解释性、模型调试与神经网络机制分析的学者

随着人工智能在多领域取得显著进展,理解模型所学表征及其概念编码方式,对科学进步与可信部署愈发重要。近期机制可解释性研究发现,神经网络常以表示空间中的线性方向表达有意义概念,并在叠加中编码多样化概念。稀疏字典学习(SDL)方法,包括稀疏自编码器、转换器与交叉编码器,通过施加稀疏约束训练辅助模型,以解耦这些叠加概念为单义特征。这些方法是现代机制可解释性的核心,但实践中普遍存在多义特征、特征吸收与死神经元现象,且缺乏理论解释。现有理论仅限于权重共享的稀疏自编码器,未能覆盖更广泛的SDL方法。本文首次构建统一理论框架,将主要SDL变体统一为分段双凸优化问题,刻画其全局解集、不可识别性及虚假极小点。该分析提供了特征吸收与死神经元的原理性解释。为在完全已知真值下暴露这些问题,我们提出线性表征基准测试(Linear Representation Bench)。基于理论指导,我们设计特征锚定技术,恢复SDL的可识别性,在合成基准与真实神经表示上显著提升特征恢复效果。

原文摘要 · Abstract (English)

As AI models achieve remarkable capabilities across diverse domains, understanding what representations they learn and how they encode concepts has become increasingly important for both scientific progress and trustworthy deployment. Recent works in mechanistic interpretability have widely reported that neural networks represent meaningful concepts as linear directions in their representation spaces and often encode diverse concepts in superposition. Various sparse dictionary learning (SDL) methods, including sparse autoencoders, transcoders, and crosscoders, are utilized to address this by training auxiliary models with sparsity constraints to disentangle these superposed concepts into monosemantic features. These methods are the backbone of modern mechanistic interpretability, yet in practice they consistently produce polysemantic features, feature absorption, and dead neurons, with very limited theoretical understanding of why these phenomena occur. Existing theoretical work is limited to tied-weight sparse autoencoders, leaving the broader family of SDL methods without formal grounding. We develop the first unified theoretical framework that casts all major SDL variants as a single piecewise biconvex optimization problem, and characterize its global solution set, non-identifiability, and spurious optima. This analysis yields principled explanations for feature absorption and dead neurons. To expose these pathologies under full ground-truth access, we introduce the Linear Representation Bench. Guided by our theory, we propose feature anchoring, a novel technique that restores SDL identifiability, substantially improving feature recovery across synthetic benchmarks and real neural representations.

可解释性稀疏学习神经机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。