通过调控视觉信息传递时机,提升多模态模型的推理准确性。
The Ebb and Flow of Multimodal Focus: Scheduling Visual Relay Windows for Grounded VLM Reasoning

- 发现多模态注意力在模型深度中经历三阶段变化,中间为关键视觉传递窗口。
- 提出TRACE框架,动态调整视觉信息传递,使接地推理平均提升4.33点,最高达6.6点。
- 适用于各类开放权重多模态模型,尤其适合依赖视觉证据的任务。
多模态视觉语言模型在基准测试中表现日益出色,但其视觉证据在进入语言处理流程后常变得不稳定,削弱了基于证据的推理能力。我们通过机制化视角分析模型内部动态,发现多模态注意力在深度上呈现稳定的三阶段重分配:早期由问题引导的组织、中期以视觉主导的传递阶段(即视觉中继窗口,VRW),以及后期返回答案构建。我们发现该中继窗口的几何结构随任务需求变化,且与接地生成存在因果关系,能区分无支撑回答与强推理路径。基于此内在节奏,我们提出TRACE——一种轻量级训练模块组成的任务自适应推理时控制框架。它在prefill阶段重塑中继分配,在解码阶段保留已整合的视觉支持。在四个开源权重的VLM骨干网络和七个基准上,TRACE在接地敏感任务中平均提升4.33点,最高达6.6点,同时改善了重推理任务的表现。结果表明,跨深度显式控制多模态焦点,是一种统一而有效的增强证据接地推理机制。
原文摘要 · Abstract (English)
Vision-language models increasingly succeed on multimodal reasoning benchmarks, yet their visual evidence often becomes unstable once it enters the language stack, weakening evidence-grounded reasoning. To understand this fragility, we examine the internal dynamics of VLMs through a mechanistic lens and uncover a stable three-stage redistribution of multimodal attention focus across depth: an early question-conditioned organization, a critical middle visual-dominant relay, and a late return to answer formation. We operationalize the middle phase as the Visual Relay Window (VRW), and show that its geometry varies with task demand, is causally tied to grounded generation, and distinguishes unsupported answers from stronger reasoning trajectories. Guided by this internal rhythm, we propose TRACE, a task-adaptive inference-time control framework with lightweight trained modules. It reshapes relay allocation during prefill and preserves assembled visual support after handoff during decoding. Across four open-weight VLM backbones and seven benchmarks, TRACE delivers large gains on grounding-sensitive settings, improving them by 4.33 points on average and by up to 6.6 points, while also improving reasoning-heavy tasks. These results show that explicitly controlling multimodal focus across depth offers a unified and effective mechanism for strengthening evidence-grounded multimodal reasoning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。