视觉语言模型推理时会逐渐失去视觉依据,导致错误自信。
Don't Blink: Evidence Collapse during Multimodal Reasoning
- 发现模型推理中对标注视觉证据的注意力下降超50%。
- 低熵但脱离视觉的预测在视觉任务中危险,符号任务中无害。
- 提出针对性视觉拦截机制,提升安全性和部署可靠性。
推理型多模态大模型在推理过程中准确率虽提升,但视觉依据逐渐丧失,形成任务相关的风险区域:低熵预测看似自信却无视觉支撑,文本监控无法察觉。在MathVista、HallusionBench和MMMU_Pro上评估三类推理型视觉语言模型,发现普遍存在的证据坍塌现象:随着推理推进,对标注证据区域的关注度显著下降,证据质量损失超过一半。全响应熵是跨数据集迁移中最可靠的纯文本不确定性信号;但仅用全局线性规则融合视觉特征时表现脆弱,常导致性能下降。熵-视觉交互模型揭示任务条件性规律:低熵且视觉脱节的预测在持续视觉参考任务中具有危害性,但在符号任务中则无害。基于此结构,设计目标性视觉否决机制,在90%覆盖率下将选择性风险降低最多1.9个百分点,且避免了预期脱节场景下的性能退化。结果支持面向任务的多模态监控,以实现分布偏移下的安全部署。
原文摘要 · Abstract (English)
Reasoning VLMs can become more accurate while progressively losing visual grounding as they think. This creates task-conditional danger zones where low-entropy predictions are confident but ungrounded, a failure mode text-only monitoring cannot detect. Evaluating three reasoning VLMs on MathVista, HallusionBench, and MMMU_Pro, we find a pervasive evidence-collapse phenomenon: attention to annotated evidence regions drops substantially, often losing over half of evidence mass, as reasoning unfolds. Full-response entropy is the most reliable text-only uncertainty signal under cross-dataset transfer, yet adding vision features with a single global linear rule is brittle and often degrades transfer. An entropy-vision interaction model reveals a task-conditional regime: lowentropy, visually disengaged predictions are hazardous on sustained visual-reference tasks but benign on symbolic tasks. Using this structure, a targeted vision veto reduces selective risk by up to 1.9 percentage points at 90% coverage, while avoiding degradations where disengagement is expected. The results support task-aware multimodal monitoring for safe deployment under distribution shift.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。