arXiv:2505.16192cs.CVcs.AI2025-05NeurIPS被引 52

让视觉语言模型学会动态聚焦图像区域,提升复杂视觉推理能力。

VLM-R$^3$: Region Recognition, Reasoning, and Refinement for Enhanced Multimodal Chain-of-Thought

  • 通过区域感知强化学习,动态决定何时、何地提取视觉证据。
  • 在MathVista等数据集上实现零样本和少样本新纪录,尤其擅长精细空间推理。
  • 适合需要高精度视觉定位与多步推理的科研或工业应用。

近期基于推理的多模态大模型在生成长序列文本推理链方面取得进展,但仍难以处理需动态反复关注和回溯视觉区域以精确对齐文本推理与视觉证据的复杂任务。我们提出VLM-R³(视觉语言模型带区域识别与推理),赋予模型三项能力:(i) 判断何时需要补充视觉证据,(ii) 精确确定图像中应关注的位置,(iii) 将相关子图像内容无缝融入交错的思维链中。核心方法为区域条件强化策略优化(R-GRPO),通过奖励模型选择信息量高的区域、制定合理变换(如裁剪、缩放)并整合视觉上下文至后续推理步骤来训练。为启动该策略,我们构建了一个精心筛选的小规模跨模态交错推理语料库(VLIR),提供逐步骤的区域选择与文本解释监督。在MathVista、ScienceQA等多个基准上的大量实验表明,VLM-R³在零样本与少样本设置下均达到新基准,尤其在依赖细微空间推理或细粒度视觉线索提取的问题上提升显著。

原文摘要 · Abstract (English)

Recently, reasoning-based MLLMs have achieved a degree of success in generating long-form textual reasoning chains. However, they still struggle with complex tasks that necessitate dynamic and iterative focusing on and revisiting of visual regions to achieve precise grounding of textual reasoning in visual evidence. We introduce \textbf{VLM-R$^3$} (\textbf{V}isual \textbf{L}anguage \textbf{M}odel with \textbf{R}egion \textbf{R}ecognition and \textbf{R}easoning), a framework that equips an MLLM with the ability to (i) decide \emph{when} additional visual evidence is needed, (ii) determine \emph{where} to ground within the image, and (iii) seamlessly weave the relevant sub-image content back into an interleaved chain-of-thought. The core of our method is \textbf{Region-Conditioned Reinforcement Policy Optimization (R-GRPO)}, a training paradigm that rewards the model for selecting informative regions, formulating appropriate transformations (e.g.\ crop, zoom), and integrating the resulting visual context into subsequent reasoning steps. To bootstrap this policy, we compile a modest but carefully curated Visuo-Lingual Interleaved Rationale (VLIR) corpus that provides step-level supervision on region selection and textual justification. Extensive experiments on MathVista, ScienceQA, and other benchmarks show that VLM-R$^3$ sets a new state of the art in zero-shot and few-shot settings, with the largest gains appearing on questions demanding subtle spatial reasoning or fine-grained visual cue extraction.

多模态推理视觉定位思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。