跨模型发现通用可解释概念,统一分析多个视觉模型内部机制
Universal Sparse Autoencoders: Interpretable Cross-Model Concept Alignment
- 用一个超完备稀疏自编码器联合学习多模型共有的概念空间
- 在多个视觉模型上发现从颜色纹理到物体部件的语义一致概念
- 适合研究多模型系统、需要可解释性分析的AI开发者
我们提出通用稀疏自编码器(USAEs),用于发现并对齐多个预训练深度神经网络中的可解释概念。与仅针对单一模型的方法不同,USAEs 联合学习一个通用概念空间,可同时重建和解释多个模型的内部激活。核心思想是训练一个单一的、过完备的稀疏自编码器(SAE),接收任意模型的激活,并解码为其他模型激活的近似表示。通过优化共享目标,学习到的字典捕捉了不同任务、架构和数据集间的共同变化因素——即通用概念。实验表明,USAEs 在多个视觉模型中发现了语义连贯且重要的通用概念,涵盖低层特征(如颜色、纹理)到高层结构(如部件、物体)。总体而言,USAEs 提供了一种强大的跨模型可解释性分析方法,并支持协同激活最大化等新应用,为多模型AI系统的深层洞察开辟新路径。
原文摘要 · Abstract (English)
We present Universal Sparse Autoencoders (USAEs), a framework for uncovering and aligning interpretable concepts spanning multiple pretrained deep neural networks. Unlike existing concept-based interpretability methods, which focus on a single model, USAEs jointly learn a universal concept space that can reconstruct and interpret the internal activations of multiple models at once. Our core insight is to train a single, overcomplete sparse autoencoder (SAE) that ingests activations from any model and decodes them to approximate the activations of any other model under consideration. By optimizing a shared objective, the learned dictionary captures common factors of variation-concepts-across different tasks, architectures, and datasets. We show that USAEs discover semantically coherent and important universal concepts across vision models; ranging from low-level features (e.g., colors and textures) to higher-level structures (e.g., parts and objects). Overall, USAEs provide a powerful new method for interpretable cross-model analysis and offers novel applications, such as coordinated activation maximization, that open avenues for deeper insights in multi-model AI systems
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。