分析视觉语言模型推理过程,发现其易受文本误导且自我修正能力有限。
Reasoning Dynamics and the Limits of Monitoring Modality Reliance in Vision-Language Models
- 追踪思维链中置信度变化,揭示模型早期判断会固化而非修正。
- 在文本主导场景下模型更易被误导,即使视觉证据充足也无法完全纠正。
- 推理型模型虽更显式引用提示词,但长篇推理可能伪装成视觉依据。
当前视觉语言模型(VLMs)具备推理能力,但其推理动态及跨模态信息整合机制尚不清晰。本研究分析了18种来自两个模型家族的指令微调与推理训练模型的推理过程,跟踪思维链(CoT)中的置信度变化,评估推理的修正效果及中间步骤贡献。结果表明,模型存在回答惯性,即早期预测承诺会被强化而非修正。尽管推理训练模型具有更强的修正能力,但其效果依赖于模态条件——从以文本为主到仅视觉场景均有差异。通过引入误导性文本线索的受控干预实验发现,模型始终受这些线索影响,即便视觉证据充分;且该影响虽可出现在思维链中,但检测难度因模型而异。推理训练模型更显式提及线索,但其流畅的思维链可能看似基于视觉却实际跟随文本,掩盖真实模态依赖。相比之下,指令微调模型提及线索较少,但较短的推理轨迹更易暴露与视觉输入的不一致。综合来看,思维链仅提供模态驱动决策的部分视图,对多模态系统的透明性与安全性有重要启示。
原文摘要 · Abstract (English)
Recent advances in vision language models (VLMs) offer reasoning capabilities, yet how these unfold and integrate visual and textual information remains unclear. We analyze reasoning dynamics in 18 VLMs covering instruction-tuned and reasoning-trained models from two different model families. We track confidence over Chain-of-Thought (CoT), measure the corrective effect of reasoning, and evaluate the contribution of intermediate reasoning steps. We find that models are prone to answer inertia, in which early commitments to a prediction are reinforced, rather than revised during reasoning steps. While reasoning-trained models show stronger corrective behavior, their gains depend on modality conditions, from text-dominant to vision-only settings. Using controlled interventions with misleading textual cues, we show that models are consistently influenced by these cues even when visual evidence is sufficient, and assess whether this influence is recoverable from CoT. Although this influence can appear in the CoT, its detectability varies across models and depends on what is being monitored. Reasoning-trained models are more likely to explicitly refer to the cues, but their longer and fluent CoTs can still appear visually grounded while actually following textual cues, obscuring modality reliance. In contrast, instruction-tuned models refer to the cues less explicitly, but their shorter traces reveal inconsistencies with the visual input. Taken together, these findings indicate that CoT provides only a partial view of how different modalities drive VLM decisions, with important implications for the transparency and safety of multimodal systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。