用强化学习让视觉语言模型学会深度反思,性能显著提升。
VL-Rethinker: Incentivizing Self-Reflection of Vision-Language Models with Reinforcement Learning

- 通过改进强化学习算法并引入反思触发机制,激发模型自我修正能力。
- 在MathVista和MathVerse上分别达到80.4%和63.5%的准确率,刷新开源模型纪录。
- 适合追求高精度多模态推理的开发者与研究者,尤其关注可解释性与慢思考系统。
近期的慢思考系统(如GPT-o1、DeepSeek-R1)在数学与科学任务中展现出显著优于快速思考模型(如GPT-4o)的能力。然而,其多模态推理表现仍与快速模型相当,例如GPT-o1在MathVista、MathVerse、MathVision等基准上的表现未有明显提升。本文提出一种无需知识蒸馏的强化学习方法,通过改进GRPO算法并引入选择性样本重放(SSR)解决优势消失问题。为进一步促进慢思考,提出强制重思机制(Forced Rethinking),在训练过程中插入反思触发符,显式要求模型进行自我验证。结合两种技术后,所提出的VL-Rethinker在MathVista、MathVerse上分别取得80.4%和63.5%的准确率,成为开源模型中的最新标杆。同时,在MathVision、MMMU-Pro、EMMA、MEGA-Bench等多个跨学科基准上也达到开源最优水平,缩小了与OpenAI-o1的差距。
原文摘要 · Abstract (English)
Recently, slow-thinking systems like GPT-o1 and DeepSeek-R1 have demonstrated great potential in solving challenging problems through explicit reflection. They significantly outperform the best fast-thinking models, such as GPT-4o, on various math and science benchmarks. However, their multimodal reasoning capabilities remain on par with fast-thinking models. For instance, GPT-o1's performance on benchmarks like MathVista, MathVerse, and MathVision is similar to fast-thinking models. In this paper, we aim to enhance the slow-thinking capabilities of vision-language models using reinforcement learning (without relying on distillation) to advance the state of the art. First, we adapt the GRPO algorithm with a novel technique called Selective Sample Replay (SSR) to address the vanishing advantages problem. While this approach yields strong performance, the resulting RL-trained models exhibit limited self-reflection or self-verification. To further encourage slow-thinking, we introduce Forced Rethinking, which appends a rethinking trigger token to the end of rollouts in RL training, explicitly enforcing a self-reflection reasoning step. By combining these two techniques, our model, VL-Rethinker, advances state-of-the-art scores on MathVista, MathVerse to achieve 80.4%, 63.5% respectively. VL-Rethinker also achieves open-source SoTA on multi-disciplinary benchmarks such as MathVision, MMMU-Pro, EMMA, and MEGA-Bench, narrowing the gap with OpenAI-o1. Our empirical results show the effectiveness of our approaches.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。