解析稀疏自编码器如何从数据中提取可解释概念的理论机制。
How Optimality Structures Sparse Dictionaries: A Theory for Understanding SAE Representations

- 基于最优性结构分析,揭示SAE特征与数据分布的内在关联。
- 解释了层级分裂、残差结构等观测现象背后的数学原理。
- 为改进SAE设计提供理论依据,适合模型可解释性研究者。
稀疏自编码器(SAEs)在解析神经表示为可解释概念方面取得成功,为理解与控制模型提供了基础。然而,SAEs究竟提取了什么,以及由此能得出哪些科学结论,并不明确。实证上,结果显而易见:SAEs学习到可解释特征。理论上,我们缺乏对‘概念’需满足何种性质才能被SAE提取的清晰解释。已有识别性研究多聚焦于简单数据生成模型(如稀疏独立特征),难以逼近互联网规模语言模型的表征。本文避开数据生成假设,直接探讨任意字典学习最优解必须满足的性质。具体地,将局部最优性分析(Gribonval & Schnass, 2010)扩展至标准SAE所近似的非负联合优化问题,推导出最优SAE特征与其分布之间的约束关系。利用这些约束,解释了一系列观测到的SAE行为:层级分裂与吸收、残差结构、密集反向特征等,均反映了L1正则与非负性如何与数据共同塑造最优字典。最后,构建了一个新颖的大字典凸优化问题,并探索了每样本原子数趋于无穷的极限情形。总体目标是将模型假设从意外观测中剥离,从而更深入理解SAE的成功之处,并为下一代SAE的设计提供原则。
原文摘要 · Abstract (English)
Sparse Autoencoders (SAEs) have found success parsing neural representations into interpretable concepts, providing a basis for understanding and control. However, what exactly SAEs extract, and, correspondingly, the scientific conclusions we can draw from them, are not obvious. Empirically, the proof is in the pudding: SAEs learn interpretable features. Theoretically, we lack a clear account of what properties a 'concept' must satisfy for an SAE to extract it. There has been extensive identifiability work studying the conditions under which sparse coding recovers ground-truth features; however, these approaches tends to focus on simple data-generating models (e.g. sparse independent features) which poorly approximate the internet-swallowing language-model representations on which SAEs are trained. Here, avoiding data-generating models, we ask simply what properties any dictionary learning optimum must satisfy. Concretely, we extend local optimality analyses (Gribonval & Schnass, 2010) to the nonnegative joint-optimisation problem that vanilla SAEs approximate, and derive constraints relating optimal SAE features to their distributions. We use these constraints to explain a range of observed SAE behaviours - hierarchical splitting & absorption, the structure of residuals, and dense antipodal features - each reflecting how L1+nonnegativity interact with data to structure optimal dictionaries. Finally, we construct a novel large-dictionary convex problem and explore the wide atom-per-datapoint limit. In sum, we hope to tease model assumptions from unexpected observations, letting us learn more from SAEs' successes and provide principles for designing their successors.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。