arXiv:2506.14002cs.LGcs.AI2025-06被引 3

提出可证明恢复语义特征的稀疏自编码器算法,提升大模型可解释性。

Taming Polysemanticity in LLMs: Provable Feature Recovery via Sparse Autoencoders

  • 将多义特征建模为单义概念的稀疏混合,定义新可识别性
  • 理论证明算法能正确恢复所有单义特征,且在15亿参数模型上表现更优
  • 适合关注大模型机制解释、可信赖AI的科研人员

我们研究使用稀疏自编码器(SAE)实现大语言模型可解释性的理论基础特征恢复问题。现有SAE训练方法缺乏严格的数学保证,且存在超参数敏感和不稳定的实践限制。为此,我们提出一种新的统计框架,将多义特征建模为底层单义概念的稀疏混合,引入特征可识别性新定义。基于此框架,提出基于「偏置自适应」的新SAE训练算法,理论上证明该算法在所提统计模型下能正确恢复所有单义特征。进一步开发了改进的实证版本——组偏置自适应(GBA),在高达1.5亿参数的LLMs上优于基准方法。本工作首次提供了具有理论恢复保证的SAE算法,是推动更透明、可信AI系统发展的关键一步。

原文摘要 · Abstract (English)

We study the challenge of achieving theoretically grounded feature recovery using Sparse Autoencoders (SAEs) for the interpretation of Large Language Models. Existing SAE training algorithms often lack rigorous mathematical guarantees and suffer from practical limitations such as hyperparameter sensitivity and instability. To address these issues, we first propose a novel statistical framework for the feature recovery problem, which includes a new notion of feature identifiability by modeling polysemantic features as sparse mixtures of underlying monosemantic concepts. Building on this framework, we introduce a new SAE training algorithm based on ``bias adaptation'', a technique that adaptively adjusts neural network bias parameters to ensure appropriate activation sparsity. We theoretically \highlight{prove that this algorithm correctly recovers all monosemantic features} when input data is sampled from our proposed statistical model. Furthermore, we develop an improved empirical variant, Group Bias Adaptation (GBA), and \highlight{demonstrate its superior performance against benchmark methods when applied to LLMs with up to 1.5 billion parameters}. This work represents a foundational step in demystifying SAE training by providing the first SAE algorithm with theoretical recovery guarantees, thereby advancing the development of more transparent and trustworthy AI systems through enhanced mechanistic interpretability.

可解释性稀疏自编码器大模型分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。