arXiv:2605.31245cs.LG2026-05被引 1

让稀疏自编码器更稳定,实现可识别的语义表示。

Toward Identifiable Sparse Autoencoders

论文配图:Toward Identifiable Sparse Autoencoders
图 1 · 摘自论文原文
  • 通过调整架构和训练流程,提升稀疏自编码器的稳定性。
  • 新模型重建误差更低,学习到的语义词典更一致。
  • 适合需要可解释神经网络表示的研究者使用。

近期,稀疏自编码器(SAEs)已成为解释和交互实际神经网络表示的有力工具。尽管这是常见的经验性共识,我们还从理论上证明了SAEs具有高度不稳定性:不同训练轮次可能产生不同的概念词典和稀疏代码。我们分析了影响现实世界SAE稳定性的模型特性,并通过最小化对架构和训练过程的修改来解决这些问题。这些改进共同带来两种版本的可识别稀疏自编码器(iSAE),是标准TopK SAE的变体,具有更低的重建误差和更高的稳定性。我们通过将SAEs与传统字典学习方法关联,从理论上解释了这一改进,表明实践中学习到的词典满足近似受限等距条件,使得对应稀疏代码近乎可识别。

原文摘要 · Abstract (English)

Recently, sparse autoencoders (SAEs) have emerged as an attractive tool for interpreting and interacting with representations in practical neural networks. While it is common empirical folklore, we also show theoretically that SAEs are highly unstable: different training runs are likely to produce different concept dictionaries and sparse codes. We characterize the model properties that hinder the stability of real-world SAEs, and address each of these problems through minimal changes to the architecture and training procedure. Together, these changes yield two versions of an \textbf{i}dentifiable SAE (iSAE), a variant of the standard TopK SAE with lower reconstruction error and improved stability. We explain this improvement theoretically by connecting SAEs with traditional dictionary learning approaches, and show that the dictionaries learned in practice satisfy an approximate restricted isometry condition, rendering the corresponding sparse codes in those models near-identifiable.

稀疏编码可解释性自编码器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。