让AI看图推理更靠谱,避免忽略小细节。
Ground-R1: Incentivizing Grounded Visual Reasoning via Reinforcement Learning
- 用新算法平衡大中小图像区域的奖励分配。
- 在多个测试中准确率和看图依据都明显提升。
- 适合需要可靠视觉推理的AI系统开发者。
大型视觉语言模型虽强大,但其预测常因缺乏对视觉证据的充分依赖而不可靠。现有‘看图思考’方法存在规模偏差:训练时大区域主导奖励,导致小但关键的视觉线索被忽略,推理时产生虚假关联。为此,我们提出Ground-R1,一种基于新型尺度相对策略优化(SRPO)的去偏框架。SRPO通过尺度感知分箱与组内/组间比较,重新校准不同尺寸区域的奖励学习,实现训练时的均衡信用分配。在通用视觉语言模型、高分辨率及视觉定位基准上的实验表明,Ground-R1显著优于标准GRPO,在响应准确性和证据锚定上均取得一致提升。
原文摘要 · Abstract (English)
Large Vision-Language Models (LVLMs) have become powerful general-purpose assistants, yet their predictions often lack reliability and interpretability due to insufficient grounding in visual evidence. The emerging thinking-with-images paradigm seeks to address this issue by explicitly anchoring reasoning to image regions. However, we empirically find that most existing methods suffer from a systematic scale-driven bias in optimization, where training rewards are dominated by large visual regions, suppressing learning from small but semantically critical evidence and leading to spurious grounding at inference time. To address this limitation, we propose Ground-R1, a de-biased thinking-with-images framework trained via a novel Scale Relative Policy Optimization (SRPO) objective that replaces standard GRPO. Specifically, our SRPO recalibrates reward learning across evidence regions of different sizes through scale-aware binning and intra-/inter-bin comparisons, enabling balanced credit assignment during training. Experimental results on general LVLM, high-resolution, and visual grounding benchmarks validate the effectiveness of Ground-R1 and show that SRPO yields consistent gains over standard GRPO in both response accuracy and evidence grounding.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。