用稀疏自编码器揭示视觉模型适配时概念的重新映射机制
Sparse autoencoders reveal selective remapping of visual concepts during adaptation
- 提出PatchSAE,从CLIP视觉变压器中提取细粒度视觉概念及空间位置
- 发现适配后性能提升主要源于原有概念的重用,而非新概念生成
- 适合关注模型可解释性与适配机制的研究者
将基础模型用于特定任务已成为构建下游应用机器学习系统的一种标准方法。然而,适配过程中发生了什么机制仍不明确。本文针对CLIP视觉变压器开发了一种新的稀疏自编码器(PatchSAE),以在细粒度层面(如形状、颜色或物体语义)提取可解释的概念及其在图像块上的空间归属。我们探究这些概念如何影响下游图像分类任务中的模型输出,并研究最新的基于提示的适配技术如何改变输入与这些概念之间的关联。尽管适配前后概念激活略有变化,但多数常见适配任务中的性能提升可由原基础模型中已存在的概念解释。本工作为训练和使用视觉变压器的稀疏自编码器提供了具体框架,并深入揭示了适配机制。
原文摘要 · Abstract (English)
Adapting foundation models for specific purposes has become a standard approach to build machine learning systems for downstream applications. Yet, it is an open question which mechanisms take place during adaptation. Here we develop a new Sparse Autoencoder (SAE) for the CLIP vision transformer, named PatchSAE, to extract interpretable concepts at granular levels (e.g., shape, color, or semantics of an object) and their patch-wise spatial attributions. We explore how these concepts influence the model output in downstream image classification tasks and investigate how recent state-of-the-art prompt-based adaptation techniques change the association of model inputs to these concepts. While activations of concepts slightly change between adapted and non-adapted models, we find that the majority of gains on common adaptation tasks can be explained with the existing concepts already present in the non-adapted foundation model. This work provides a concrete framework to train and use SAEs for Vision Transformers and provides insights into explaining adaptation mechanisms.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。