arXiv:2605.14694cs.LG2026-05被引 1

揭示稀疏自编码器中解释性、效率与重构精度的内在权衡

The Rate-Distortion-Polysemanticity Tradeoff in SAEs

论文配图:The Rate-Distortion-Polysemanticity Tradeoff in SAEs
图 1 · 摘自论文原文
  • 提出率-失真-多语义性三重权衡理论,解释为何高可解释性难实现
  • 证明强制单语义会导致编码率和重构误差上升,数据分布决定最优多语义程度
  • 验证真实语言模型中现有度量的有效性,强调需从数据出发设计模型

稀疏自编码器(SAEs)在准确重构输入(低失真)、高效使用少量特征(低率)的同时,常难以学习到单语义表示(高可解释性),限制了其在机制可解释性中的应用。本文首次刻画了学习忠实、高效且可解释表征之间的张力,提出SAE中的率-失真-多语义性权衡。在理想建模假设下,我们理论与实证表明,强制单语义必然导致率与失真增加。基于输入观测背后的生成模型假设,进一步证明最优SAE的多语义程度由训练数据分布决定,尤其取决于特征共现概率。最后,我们在真实场景下拓展分析,推导出当数据生成过程未知时,多语义性度量应满足的必要条件,并对在大语言模型上训练的SAE进行了现有代理度量的基准测试。结果表明,多语义性本质上是数据问题,必须在架构与优化层面予以考虑。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) that can accurately reconstruct their input (minimizing distortion) by making efficient use of few features (minimizing the rate) often fail to learn monosemantic representations (highly interpretable), limiting their usefulness for mechanistic interpretability. In this paper, we characterise this tension in learning faithful, efficient, and interpretable explanations, introducing the Rate-Distortion-Polysemanticity tradeoff in SAEs. Under toy-modeling assumptions, we theoretically and empirically show that restricting the SAE to be monosemantic necessarily comes with an increase in rate and distortion. Assuming a generative model behind the input observations, we further demonstrate that the degree of polysemanticity of optimal SAEs is determined by the training data distribution, especially by the probability of features to co-occur. Finally, we extend the analysis to real-world settings by deriving necessary conditions that a polysemanticity measure should satisfy when the data-generating process is unknown, and we benchmark existing proxy metrics on SAEs trained on Large Language Models. Taken together, our findings show that polysemanticity is a data problem that should be accounted for when addressing it at the architectural and optimization level.

稀疏编码可解释性语言模型机器学习理论

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。