让稀疏自编码器能主动探测用户定义的概念,实现可控且可逆的语义接口。
Concept-SAE: A Controllable and Invertible Concept Interface for Sparse Autoencoders
- 通过双监督将激活空间分解为概念令牌与自由令牌,实现语义对齐与空间定位。
- 在概念检测、可控编辑和抗干扰测试中均优于现有方法,表现更稳定可靠。
- 适合需要精准诊断模型内部概念的科研人员和可解释性研究者使用。
标准稀疏自编码器(SAE)擅长发现模型学习到的特征字典,为被动特征发现提供了有力视角。但其被动特性使系统性评估或分析用户关注的概念变得困难。我们提出Concept-SAE,一种为SAE引入结构化可控接口的框架,以探查用户定义的概念。Concept-SAE将激活子空间分解为两个正交分量:概念令牌通过概念存在性和空间定位双重监督与外部语义对齐;自由令牌则像标准SAE一样捕捉剩余信息。这种混合解耦策略确保概念令牌忠实、空间定位准确,且与残余子空间清晰分离,同时保留SAE开放发现新概念的能力。大量实验表明,Concept-SAE生成的表示具有高保真度、良好定位性和强解耦性,在接口质量上超越替代方案。最后,我们通过三项诊断评估验证其效用:对抗样本分类检测、可控反事实编辑控制性测试、以及对抗扰动下的稳定性测试。结果共同表明,Concept-SAE为SAE提供了可靠的机制,用于评估、探查和诊断用户定义的概念。
原文摘要 · Abstract (English)
Standard Sparse Autoencoders (SAEs) excel at discovering a dictionary of a model's learned features, providing a powerful lens for passive feature discovery. However, this passive nature makes it difficult to systematically evaluate or analyze concepts that users explicitly care about. We introduce Concept-SAE, a framework that augments SAEs with a structured and controllable interface for probing user-defined concepts. Concept-SAE decomposes an activation subspace into two orthogonal components: Concept Tokens, which are aligned to externally specified semantics through dual supervision on both concept existence and spatial localization, and Free Tokens, which operate like standard SAEs to capture all remaining information. This hybrid disentanglement strategy ensures that Concept Tokens are faithful, spatially grounded, and cleanly separated from the residual subspace while preserving the ability of SAEs for open-ended concept discovery. We conduct extensive experiments demonstrating that Concept-SAE yields high-fidelity, well-localized, and strongly disentangled concept representations, outperforming alternatives in interface quality. Finally, we validate the utility of this conceptual interface through three diagnostic evaluations: a detection test on classifying adversarial image samples, a controllability test focusing on controlled counterfactual editing and a stability test using adversarial perturbations. Together, these results show that Concept-SAE equips SAEs with a reliable mechanism for evaluating, probing, and diagnosing user-defined concepts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。