arXiv:2508.11277cs.CVcs.LG2025-08ICCV被引 9

用稀疏自编码器挖掘视觉模型内部语义,提升可解释性与可控生成。

Probing the Representational Power of Sparse Autoencoders in Vision Models

  • 在视觉模型中使用稀疏自编码器学习可解释特征。
  • 特征能提升分布外检测与跨模型泛化能力。
  • 适用于图像嵌入、多模态模型和扩散模型,适合研究可解释性者。

稀疏自编码器(SAEs)已成为解析大语言模型隐藏状态的流行工具。通过从稀疏瓶颈层重建激活值,SAEs能够从语言模型的高维内部表示中发现语义明确的特征。尽管在语言模型中广受关注,其在视觉领域的研究仍较薄弱。本文对三类视觉模型架构——图像嵌入模型、多模态大模型(LMMs)和扩散模型——进行了广泛的评估,验证了SAEs在视觉任务中的表征能力。实验表明,SAE特征具有语义意义,能增强分布外(OOD)泛化能力,并支持可控生成。在图像嵌入模型中,学到的SAE特征可用于OOD检测,并揭示模型底层的本体结构;在扩散模型中,可通过操控文本编码器实现语义控制,并构建自动化流程发现人类可读属性;初步探索多模态大模型发现,SAE特征揭示了视觉与语言模态间的共享表示。本研究为视觉模型中SAE评估提供了基础,凸显其在提升可解释性、泛化性和可控性方面的巨大潜力。

原文摘要 · Abstract (English)

Sparse Autoencoders (SAEs) have emerged as a popular tool for interpreting the hidden states of large language models (LLMs). By learning to reconstruct activations from a sparse bottleneck layer, SAEs discover interpretable features from the high-dimensional internal representations of LLMs. Despite their popularity with language models, SAEs remain understudied in the visual domain. In this work, we provide an extensive evaluation the representational power of SAEs for vision models using a broad range of image-based tasks. Our experimental results demonstrate that SAE features are semantically meaningful, improve out-of-distribution generalization, and enable controllable generation across three vision model architectures: vision embedding models, multi-modal LMMs and diffusion models. In vision embedding models, we find that learned SAE features can be used for OOD detection and provide evidence that they recover the ontological structure of the underlying model. For diffusion models, we demonstrate that SAEs enable semantic steering through text encoder manipulation and develop an automated pipeline for discovering human-interpretable attributes. Finally, we conduct exploratory experiments on multi-modal LLMs, finding evidence that SAE features reveal shared representations across vision and language modalities. Our study provides a foundation for SAE evaluation in vision models, highlighting their strong potential improving interpretability, generalization, and steerability in the visual domain.

视觉模型可解释性稀疏编码可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。