首次实现视觉语言模型内部机制的透明追踪,揭示多模态推理路径。
Circuit Tracing in Vision-Language Models: Understanding the Internal Mechanisms of Multimodal Thinking
- 通过编码器、归因图与注意力方法,系统追踪多模态融合路径。
- 发现不同视觉特征通路可支持数学推理与跨模态关联。
- 验证通路具因果性与可控性,适用于可解释性研究者。
视觉语言模型(VLMs)虽强大却仍为黑箱。本文提出首个透明电路追踪框架,系统分析多模态推理机制。利用转码器、归因图与基于注意力的方法,揭示了视觉与语义概念的分层整合过程。研究发现,不同视觉特征通路可分别处理数学推理并支持跨模态关联。通过特征引导与电路修补验证,证明这些通路具有因果性与可控性,为构建更可解释、更可靠的VLM奠定基础。
原文摘要 · Abstract (English)
Vision-language models (VLMs) are powerful but remain opaque black boxes. We introduce the first framework for transparent circuit tracing in VLMs to systematically analyze multimodal reasoning. By utilizing transcoders, attribution graphs, and attention-based methods, we uncover how VLMs hierarchically integrate visual and semantic concepts. We reveal that distinct visual feature circuits can handle mathematical reasoning and support cross-modal associations. Validated through feature steering and circuit patching, our framework proves these circuits are causal and controllable, laying the groundwork for more explainable and reliable VLMs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。