arXiv:2604.28119cs.LGcs.AI2026-04被引 16

揭示稀疏自编码器如何捕捉概念流形,指出其存在局部与全局两种机制。

Do Sparse Autoencoders Capture Concept Manifolds?

论文配图:Do Sparse Autoencoders Capture Concept Manifolds?
图 1 · 摘自论文原文
  • 提出理论框架,区分稀疏自编码器对流形的全局覆盖与局部镶嵌两种方式。
  • 实证发现现有架构仅部分恢复连续结构,呈现碎片化稀释现象。
  • 建议未来方法以几何对象为可解释单元,而非单一线性方向。

稀疏自编码器(SAEs)被广泛用于从神经网络表征中提取可解释特征,常基于概念对应独立线性方向的隐含假设。然而,越来越多证据表明,许多概念沿低维流形组织,编码连续几何关系。本文提出三个核心问题:何为捕捉流形?现有SAE架构在何时能实现?如何实现?我们构建理论框架回答这些问题,发现SAEs可通过两种根本不同的方式捕捉流形:全局方式,即分配一组紧凑的基元,其线性张成包含整个流形;或局部方式,将流形分布于多个特征上,各自选择性地覆盖几何的特定区域。实证结果表明,当前SAE仅次优地恢复连续结构,混合了全局子空间与局部镶嵌解,处于一种被称为‘稀释’的碎片化状态。这解释了为何流形结构通常难以在单个概念层面显现,并推动事后无监督发现方法转向寻找原子组的相干结构,而非孤立方向。更广泛地,我们的研究提示,未来的表征学习应将几何对象而非单一方向视为可解释性的基本单元。

原文摘要 · Abstract (English)

Sparse autoencoders (SAEs) are widely used to extract interpretable features from neural network representations, often under the implicit assumption that concepts correspond to independent linear directions. However, a growing body of evidence suggests that many concepts are instead organized along low-dimensional manifolds encoding continuous geometric relationships. This raises three basic questions: what does it mean for an SAE to capture a manifold, when do existing SAE architectures do so, and how? We develop a theoretical framework that answers these questions and show that SAEs can capture manifolds in two fundamentally different ways: globally, by allocating a compact group of atoms whose linear span contains the entire manifold, or locally, by distributing it across features that each selectively tile a restricted region of the underlying geometry. Empirically, we find that SAEs suboptimally recover continuous structures, mixing the global subspace and local tiling solutions in a fragmented regime we call dilution. This explains why manifold structure is rarely visible at the level of individual concepts and motivates post-hoc unsupervised discovery methods that search for coherent groups of atoms rather than isolated directions. More broadly, our results suggest that future representation learning methods should treat geometric objects, not just individual directions, as the basic units of interpretability.

稀疏自编码器概念流形可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。