arXiv:2602.03733cs.CV2026-02中稿 · ICLR被引 3

让AI看图推理时能一步步聚焦区域,更准更可信。

RegionReasoner: Region-Grounded Multi-Round Visual Reasoning

  • 用强化学习强制每步推理必须引用具体图像区域
  • 多轮推理准确率提升,空间定位更精确
  • 适合研究视觉推理与可解释AI的学者

大型视觉语言模型在视觉推理上已取得显著进展,但现有系统多依赖单步或纯文本推理,难以在多个视觉上下文中迭代优化理解。为此,我们引入一个涵盖检测与分割任务的新多轮视觉推理基准,支持系统性评估。同时提出RegionReasoner,一种基于强化学习的框架,通过要求每个推理步骤明确引用对应参考边界框来实现视觉接地,并利用全局-局部一致性奖励维持语义连贯性。该奖励从整体场景描述和区域级描述中提取关键对象与名词,与推理过程对齐以确保跨步骤一致性。RegionReasoner采用结合接地保真度与全局-局部语义对齐的结构化奖励进行优化。在检测与分割任务上的实验表明,RegionReasoner-7B配合新基准RegionDial-Bench,在多轮推理准确率、空间定位精度及全局-局部一致性方面均有显著提升,为该新兴方向建立了强基线。

原文摘要 · Abstract (English)

Large vision-language models have achieved remarkable progress in visual reasoning, yet most existing systems rely on single-step or text-only reasoning, limiting their ability to iteratively refine understanding across multiple visual contexts. To address this limitation, we introduce a new multi-round visual reasoning benchmark with training and test sets spanning both detection and segmentation tasks, enabling systematic evaluation under iterative reasoning scenarios. We further propose RegionReasoner, a reinforcement learning framework that enforces grounded reasoning by requiring each reasoning trace to explicitly cite the corresponding reference bounding boxes, while maintaining semantic coherence via a global-local consistency reward. This reward extracts key objects and nouns from both global scene captions and region-level captions, aligning them with the reasoning trace to ensure consistency across reasoning steps. RegionReasoner is optimized with structured rewards combining grounding fidelity and global-local semantic alignment. Experiments on detection and segmentation tasks show that RegionReasoner-7B, together with our newly introduced benchmark RegionDial-Bench, considerably improves multi-round reasoning accuracy, spatial grounding precision, and global-local consistency, establishing a strong baseline for this emerging research direction.

视觉推理多轮对话强化学习图像定位

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。