arXiv:2502.12892cs.CV2025-02ICML被引 71

让视觉模型的解释更稳定,新方法能可靠提取语义概念。

Archetypal SAE: Adaptive and Stable Dictionary Learning for Concept Extraction in Large Vision Models

  • 用凸包约束字典原子,提升概念提取的稳定性
  • 在多个基准上优于现有方法,能发现新语义概念
  • 适合关注模型可解释性的研究者与工程师

稀疏自编码器(SAEs)是机器学习可解释性的重要工具,能无监督地将模型表征分解为一组抽象、可理解的概念字典。然而我们发现,现有SAEs存在严重不稳定性:相同模型在相似数据上训练却生成差异显著的字典,削弱其可靠性。为此,我们受Cutler & Breiman(1994)提出的原型分析启发,提出弧形原型SAE(A-SAE),将字典原子限制在数据的凸包内。这种几何约束显著提升了字典稳定性,其适度放松的变体RA-SAE在重建能力上达到当前最优水平。为严格评估字典质量,我们引入两个新基准:(i)合理性——是否恢复“真实”分类方向;(ii)可辨识性——是否解耦合成概念混合。在所有评测中,RA-SAE始终生成更结构化的表示,并在大规模视觉模型中揭示了新的语义概念。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have emerged as a powerful framework for machine learning interpretability, enabling the unsupervised decomposition of model representations into a dictionary of abstract, human-interpretable concepts. However, we reveal a fundamental limitation: existing SAEs exhibit severe instability, as identical models trained on similar datasets can produce sharply different dictionaries, undermining their reliability as an interpretability tool. To address this issue, we draw inspiration from the Archetypal Analysis framework introduced by Cutler & Breiman (1994) and present Archetypal SAEs (A-SAE), wherein dictionary atoms are constrained to the convex hull of data. This geometric anchoring significantly enhances the stability of inferred dictionaries, and their mildly relaxed variants RA-SAEs further match state-of-the-art reconstruction abilities. To rigorously assess dictionary quality learned by SAEs, we introduce two new benchmarks that test (i) plausibility, if dictionaries recover "true" classification directions and (ii) identifiability, if dictionaries disentangle synthetic concept mixtures. Across all evaluations, RA-SAEs consistently yield more structured representations while uncovering novel, semantically meaningful concepts in large-scale vision models.

可解释性视觉模型稀疏编码概念提取

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。