arXiv:2606.02061cs.LG2026-06被引 1

Archetypal SAE 的稳定性是初始化和度量设计的产物,非真实改进。

Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design

论文配图:Ablating Archetypes: The Stability of Archetypal SAEs is an Artifact of Initialization and Metric Design
图 1 · 摘自论文原文
  • 通过固定初始值模拟稳定性,实则掩盖了训练过程中的随机性差异。
  • 移除固定初始化后,其稳定优势消失,表明并非真正提升模型鲁棒性。
  • 强调需结合轨迹分析与初始化消融,才能真实评估特征可复现性。

稀疏自编码器(SAE)在神经网络激活上进行字典学习,生成可解释且减少多义性的过完备基。然而,不同随机种子下的SAE特征存在显著差异,即不稳定性问题。为此,Fel等人(2025)提出谱型SAE(Archetypal SAE),声称能更可靠地提取概念,并报告训练结束时具有更高的字典稳定性。我们发现,这种稳定性实为多轮实验中使用相同初始化所致。通过分析,我们厘清了机制可解释性中的两个关键概念:稳定性指独立训练模型间的共识,而稳定化是多个独立初始化趋向同一解的过程。此区分对自然语言处理中的机制可解释性至关重要,因特征稳定性常被当作可复用分析单元的证据。实验显示,谱型SAE共享确定性的k-means解码器初始化,使运行间字典距离从零开始;一旦移除该初始化,谱型约束不再提供稳定化优势。此外,我们识别出预处理依赖的余弦几何问题,干扰终点稳定性度量的解读。总体而言,本研究支持将SAE置于传统字典学习框架中考察,同时指出稳定性声明需依赖轨迹诊断与初始化消融验证。

原文摘要 · Abstract (English)

Dictionary learning with sparse autoencoders (SAEs) produces overcomplete bases from neural network activations that are often interpretable and reduces polysemanticity. However, features from SAEs vary substantially across random seeds -- a problem known as instability. Archetypal SAEs (Fel et al., 2025) were proposed as a general dictionary-learning intervention for more reliable concept extraction, and report more stable dictionaries at the end of training. We demonstrate that the stability claimed by archetypal SAEs is a result of setting identical initialization across multiple runs. Through our analyses, we attempt to clarify two distinct notions in mechanistic interpretability that may be ambiguously used: stability is agreement between two independently trained models, whereas stabilization is the convergence of independently initialized runs toward a common solution. This distinction is critical for mechanistic interpretability of natural language processing (NLP), where feature stability is increasingly used as evidence that SAE features are reusable units of analysis. Experiments from archetypal SAEs share a deterministic k-means decoder initialization, setting inter-run dictionary distance to zero before training begins. When this initialization is removed, the archetypal constraint provides no stabilization advantage in our setting. We further identify a preprocessing-dependent cosine geometry issue that complicates interpretation of endpoint stability metrics. Overall, our study supports the value of studying SAEs within the larger dictionary-learning tradition while showing that stability claims require trajectory diagnostics and initialization ablations.

SAE稳定性可解释性初始化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。