arXiv:2603.22593cs.CVcs.AI2026-03中稿 · CVPR

用语言模型解释视觉特征,无需人工干预。

Language Models Can Explain Visual Features via Steering

  • 通过控制视觉编码器的稀疏自编码器特征,让语言模型描述其‘看到’的内容。
  • 解释质量随语言模型规模增大而提升,且方法可扩展。
  • 适合关注模型可解释性、视觉表征分析的研究者。

稀疏自编码器(Sparse Autoencoders, SAE)在视觉模型中揭示了数千个特征,但如何在无需人工干预的情况下解释这些特征仍是开放问题。以往方法依赖激活最高的输入样本生成相关性解释,本文提出一种基于因果干预的新思路:利用视觉-语言模型结构,在输入为空图像时对视觉编码器中的单个SAE特征进行控制,随后提示语言模型描述其‘所见’,从而获取该特征对应的视觉概念。结果表明,该方法具有可扩展性,可补充传统的基于输入样本的解释方式,为视觉模型的自动化可解释性提供新视角。此外,解释质量随语言模型规模增加而持续提升,凸显该方法的潜力。最后,我们提出Steering-informed Top-k,融合因果干预与输入基方法的优势,实现最佳解释质量且无额外计算开销。

原文摘要 · Abstract (English)

Sparse Autoencoders uncover thousands of features in vision models, yet explaining these features without requiring human intervention remains an open challenge. While previous work has proposed generating correlation-based explanations based on top activating input examples, we present a fundamentally different alternative based on causal interventions. We leverage the structure of Vision-Language Models and steer individual SAE features in the vision encoder after providing an empty image. Then, we prompt the language model to explain what it ``sees'', effectively eliciting the visual concept represented by each feature. Results show that Steering offers an scalable alternative that complements traditional approaches based on input examples, serving as a new axis for automated interpretability in vision models. Moreover, the quality of explanations improves consistently with the scale of the language model, highlighting our method as a promising direction for future research. Finally, we propose Steering-informed Top-k, a hybrid approach that combines the strengths of causal interventions and input-based approaches to achieve state-of-the-art explanation quality without additional computational cost.

可解释性视觉语言模型特征编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。