用强化学习提升视觉理解能力,首次在COCO上达到31.9% AP
Perception-R1: Pioneering Perception Policy with Reinforcement Learning

- 基于GRPO的强化学习框架,用于多模态模型后训练中的感知策略学习
- 在多个视觉任务上显著提效,最高提升17.9%,首次实现COCO 31.9% AP
- 揭示感知复杂度与奖励设计对强化学习效果的关键影响,适合视觉理解研究者
受DeepSeek-R1成功的启发,我们探索了基于规则的强化学习(RL)在多模态大模型(MLLM)后训练中感知策略学习的潜力。尽管前景可观,但初步实验表明,通过强化学习引入思维过程,并未在所有视觉感知任务中稳定提升性能。这促使我们回归基础,探究强化学习在不同感知任务中的实际作用。我们发现,感知复杂度是决定强化学习有效性的重要因素,且奖励设计对逼近模型感知上限至关重要。基于此,我们提出Perception-R1,一个在MLLM后训练中使用GRPO的可扩展强化学习框架。以标准Qwen2.5-VL-3B-Instruct为基础,Perception-R1在RefCOCO+上提升4.2%,在PixMo-Count上提升17.9%,在PageOCR上提升4.2%,并在COCO2017 val上首次达到31.9% AP,为感知策略学习建立了强基准。
原文摘要 · Abstract (English)
Inspired by the success of DeepSeek-R1, we explore the potential of rule-based reinforcement learning (RL) in MLLM post-training for perception policy learning. While promising, our initial experiments reveal that incorporating a thinking process through RL does not consistently lead to performance gains across all visual perception tasks. This leads us to delve into the essential role of RL in the context of visual perception. In this work, we return to the fundamentals and explore the effects of RL on different perception tasks. We observe that the perceptual complexity is a major factor in determining the effectiveness of RL. We also observe that reward design plays a crucial role in further approching the upper limit of model perception. To leverage these findings, we propose Perception-R1, a scalable RL framework using GRPO during MLLM post-training. With a standard Qwen2.5-VL-3B-Instruct, Perception-R1 achieves +4.2% on RefCOCO+, +17.9% on PixMo-Count, +4.2% on PageOCR, and notably, 31.9% AP on COCO2017 val for the first time, establishing a strong baseline for perception policy learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。