让多模态大模型学会自我反思,提升复杂问题推理能力。
SRPO: Enhancing Multimodal LLM Reasoning via Reflection-Aware Reinforcement Learning
- 用高级多模态模型生成反思数据,训练模型学会自我修正。
- 在多个基准上,准确率显著超越现有模型,最高提升12.3%。
- 适合研究多模态推理、自省机制的学者和开发者。
多模态大语言模型在推理任务中展现出潜力,但在需要显式自我反思与修正的复杂问题上仍表现不足,尤其相比单模态文本模型。现有反思方法过于简单,难以生成有意义的反馈,因预训练模型的推理能力与知识在初始训练后基本固定。为此,我们提出基于组相对策略优化的多模态自我反思增强框架(SRPO),分两阶段提升多模态大模型的推理能力。第一阶段,借助先进多模态模型生成高质量、聚焦反思的数据集,指导策略模型学习推理与自我反思。第二阶段,在GRPO框架中引入新型奖励机制,鼓励简洁且认知上有意义的反思,避免冗余。在MathVista、MathVision、MathVerse和MMMU-Pro等多个多模态推理基准上,使用Qwen-2.5-VL-7B和Qwen-2.5-VL-32B进行实验,结果表明SRPO显著优于当前最优模型,在推理准确率与反思质量上均有明显提升。
原文摘要 · Abstract (English)
Multimodal large language models (MLLMs) have shown promising capabilities in reasoning tasks, yet still struggle with complex problems requiring explicit self-reflection and self-correction, especially compared to their unimodal text-based counterparts. Existing reflection methods are simplistic and struggle to generate meaningful and instructive feedback, as the reasoning ability and knowledge limits of pre-trained models are largely fixed during initial training. To overcome these challenges, we propose Multimodal Self-Reflection enhanced reasoning with Group Relative Policy Optimization (SRPO), a two-stage reflection-aware reinforcement learning (RL) framework explicitly designed to enhance multimodal LLM reasoning. In the first stage, we construct a high-quality, reflection-focused dataset under the guidance of an advanced MLLM, which generates reflections based on initial responses to help the policy model learn both reasoning and self-reflection. In the second stage, we introduce a novel reward mechanism within the GRPO framework that encourages concise and cognitively meaningful reflection while avoiding redundancy. Extensive experiments across multiple multimodal reasoning benchmarks, including MathVista, MathVision, MathVerse, and MMMU-Pro, using Qwen-2.5-VL-7B and Qwen-2.5-VL-32B demonstrate that SRPO significantly outperforms state-of-the-art models, achieving notable improvements in both reasoning accuracy and reflection quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。