用外部模型生成高质量动作,提升视觉推理模型的性能与训练效率
Vision-EKIPL: External Knowledge-Infused Policy Learning for Visual Reasoning
- 引入外部辅助模型生成高质动作,扩展策略模型探索空间
- 在Reason-RFT-CoT上实现最高5%的性能提升,加速收敛
- 适合追求视觉推理能力突破的研究者与应用开发者
视觉推理对理解复杂多模态数据、推动通用人工智能发展至关重要。现有方法通过强化学习微调(如GRPO)提升多模态大语言模型(MLLMs)的推理能力,但当前强化学习方法仅从策略模型自身采样动作组,限制了推理上限并导致训练效率低下。为此,本文提出一种新型强化学习框架Vision-EKIPL,其核心是在强化学习训练过程中引入由外部辅助模型生成的高质量动作,指导策略模型优化。知识注入使模型探索空间显著扩大,有效提升推理边界,大幅加快训练收敛速度与效率。实验结果表明,Vision-EKIPL在Reason-RFT-CoT基准上相比最先进方法性能提升最高达5%,验证了该方法可克服传统强化学习局限,显著增强MLLMs的视觉推理能力,并为该领域研究提供新范式。
原文摘要 · Abstract (English)
Visual reasoning is crucial for understanding complex multimodal data and advancing Artificial General Intelligence. Existing methods enhance the reasoning capability of Multimodal Large Language Models (MLLMs) through Reinforcement Learning (RL) fine-tuning (e.g., GRPO). However, current RL approaches sample action groups solely from the policy model itself, which limits the upper boundary of the model's reasoning capability and leads to inefficient training. To address these limitations, this paper proposes a novel RL framework called \textbf{Vision-EKIPL}. The core of this framework lies in introducing high-quality actions generated by external auxiliary models during the RL training process to guide the optimization of the policy model. The policy learning with knowledge infusion from external models significantly expands the model's exploration space, effectively improves the reasoning boundary, and substantially accelerates training convergence speed and efficiency. Experimental results demonstrate that our proposed Vision-EKIPL achieved up to a 5\% performance improvement on the Reason-RFT-CoT Benchmark compared to the state-of-the-art (SOTA). It reveals that Vision-EKIPL can overcome the limitations of traditional RL methods, significantly enhance the visual reasoning performance of MLLMs, and provide a new effective paradigm for research in this field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。