提出轻量级插件ROVER,高效整合多图视觉证据进行推理
ROVER: Routing Object-Centric Visual Evidence for Grounded Multi-Image Reasoning

- 基于物体中心注意力动态聚合跨图像视觉证据
- 在MM-GCoT上提升4.8%答案准确率,VideoEspresso上提升8.6%
- 无需复杂监督,适合需要多图推理的视觉语言模型
多模态大语言模型日益依赖定位和交错的视觉证据进行推理。现有基于定位的方法通常通过注入裁剪图像块或区域特定特征来聚焦感兴趣区域(RoIs),但此类设计会削弱整体场景理解与对象间关系,且解码开销随区域数量和大小增加。另一种自适应特征选择常需细粒度标注或复杂启发式规则。为此,我们提出ROVER(面向接地多图推理的物体中心视觉证据路由),一种轻量、可学习的插件,用于高效全局视觉证据路由。每次物体定位预测后,ROVER注入一个步骤特异性标记三元组,协同实现:(i) 聚合当前推理上下文,(ii) 通过物体中心差分注意力将图像内线索提炼至视觉工作空间,(iii) 在该空间中路由并整合跨对象与跨图像的历史感知证据以供后续推理。我们将ROVER集成至Qwen2.5-VL-7B,并开发交错SFT-to-GRPO训练流程。严格遵循原数据集与评估协议,方法在MM-GCoT上实现+4.8%答案准确率、+14.6%定位准确率,在VideoEspresso上实现+8.6%答案准确率。经VideoEspresso训练的模型展现出强泛化能力,在多个基准测试上平均超越基线+4.7%。
原文摘要 · Abstract (English)
Multimodal Large Language Models (MLLMs) have increasingly localized and interleaved visual evidence for deliberative reasoning. Grounding-based approaches typically focus on regions of interest (RoIs) by injecting cropped image patches or RoI-specific features into the reasoning context. However, such designs can weaken holistic scene understanding and inter-object relations, while incurring decoding costs that scale with the number and size of RoIs. Alternatively, adaptive visual feature selection often requires fine-grained supervision or complex heuristics. To address these limitations, we propose ROVER (Routing Object-centric Visual Evidence for grounded multi-image Reasoning), a lightweight, learnable plugin for efficient global visual evidence routing. Upon each object grounding prediction, ROVER injects a step-specific token triplet to synergistically: (i) aggregate the ongoing reasoning context, (ii) distill intra-image cues into a visual working space via object-centric differential attention, and (iii) route and integrate history-aware evidence across objects and images within this space for subsequent reasoning. We integrate ROVER into Qwen2.5-VL-7B and develop an interleaved SFT-to-GRPO training pipeline. Strictly adhering to the original datasets and evaluation protocols, our method achieves the best performance on MM-GCoT (+4.8% answer accuracy, +14.6% grounding accuracy) and VideoEspresso (+8.6% answer accuracy). The VideoEspresso-trained model demonstrates strong transferability, outperforming the base model by +4.7% on average across diverse benchmarks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。