arXiv:2503.01822cs.LGcs.AI2025-03NeurIPS被引 58

不同稀疏自编码器会因结构假设差异,导致发现的概念完全不同。

Projecting Assumptions: The Duality Between Sparse Autoencoders and Concept Geometry

  • 将SAE建模为双层优化问题,揭示其编码假设对概念探测的决定性影响
  • 在真实模型激活上验证:忽略概念维度与非线性分离性会导致漏检
  • 提出新SAE架构,可发现此前隐藏的概念,适合注重解释力的开发者

稀疏自编码器(SAE)被广泛用于从神经网络表征中识别有意义的概念。然而,它们是否真正揭示了模型依赖的所有概念,还是存在固有偏见?本文提出统一框架,将SAE视为双层优化问题的解,揭示其根本挑战:每个SAE都对概念在表征中的编码方式施加结构假设,进而决定其能检测或无法检测的概念。这意味着不同SAE不可互换——更换架构可能暴露全新概念,也可能掩盖已有概念。我们通过三类实验系统探究此效应:控制玩具模型、半合成真实激活实验,以及大规模自然数据集。研究考察现实概念常具有的两个特性:内在维度异质性(部分概念本质低维,部分不是)和非线性可分性。结果表明,若忽略这些特性,SAE将无法恢复概念;我们设计的新型SAE显式融合两者,成功发现此前隐藏的概念,验证理论洞见。研究挑战了通用SAE的观念,强调可解释性需根据架构选择。总之,SAE不仅揭示概念,更决定了什么可以被看见。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) are widely used to interpret neural networks by identifying meaningful concepts from their representations. However, do SAEs truly uncover all concepts a model relies on, or are they inherently biased toward certain kinds of concepts? We introduce a unified framework that recasts SAEs as solutions to a bilevel optimization problem, revealing a fundamental challenge: each SAE imposes structural assumptions about how concepts are encoded in model representations, which in turn shapes what it can and cannot detect. This means different SAEs are not interchangeable -- switching architectures can expose entirely new concepts or obscure existing ones. To systematically probe this effect, we evaluate SAEs across a spectrum of settings: from controlled toy models that isolate key variables, to semi-synthetic experiments on real model activations and finally to large-scale, naturalistic datasets. Across this progression, we examine two fundamental properties that real-world concepts often exhibit: heterogeneity in intrinsic dimensionality (some concepts are inherently low-dimensional, others are not) and nonlinear separability. We show that SAEs fail to recover concepts when these properties are ignored, and we design a new SAE that explicitly incorporates both, enabling the discovery of previously hidden concepts and reinforcing our theoretical insights. Our findings challenge the idea of a universal SAE and underscores the need for architecture-specific choices in model interpretability. Overall, we argue an SAE does not just reveal concepts -- it determines what can be seen at all.

稀疏自编码器模型可解释性概念发现表征几何

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。