通过强化学习视角优化多模态推理的自进化训练,突破性能瓶颈。
Diving into Self-Evolving Training for Multimodal Reasoning
- 借鉴强化学习,设计三要素协同的自进化训练框架
- 提出自动平衡机制,显著缓解模型性能饱和问题
- 适用于不同规模模型与多种评测基准,效果持续提升
自进化训练——模型从自身输出中迭代学习——已成为复杂推理任务的关键方法,可缓解高质量思维链数据稀缺的问题。然而,该方法在更复杂的多模态推理领域中的有效性仍不明确,对其关键影响因素的理解也有限。此外,该训练范式面临性能饱和的核心挑战,阻碍进一步提升与扩展。受强化学习启发,本文从强化学习视角重构多模态推理的自进化训练,识别出三个关键因素:训练方法、奖励模型与提示变化。通过系统分析,确立了相对最优的设计原则,显著增强多模态推理能力。深入探究训练动态后,揭示了饱和的根本原因,并提出一种新的自动平衡机制以缓解此问题。基于上述发现,本文提出 M-STAR(Multimodal Self-evolving Training for Reasoning)框架,在不同规模模型与多样基准上均实现持续性能提升。所有资源已公开于 https://mstar-lmm.github.io。
原文摘要 · Abstract (English)
Self-evolving trainin--where models iteratively learn from their own outputs--has emerged as a key approach for complex reasoning tasks, addressing the scarcity of high-quality chain-of-thought data. However, its effectiveness in multimodal reasoning, a domain more intricate than text-only reasoning, remains underexplored, and the understanding of critical factors in this training paradigm remains limited. Furthermore, a central challenge for this training method is performance saturation, which impedes further improvements and scalability. Inspired by reinforcement learning (RL), in this paper, we reframe self-evolving training for multimodal reasoning through the lens of RL, identifying three pivotal factors: Training Method, Reward Model, and Prompt Variation. Through systematic analysis, we establish relatively optimal design principles that significantly enhance multimodal reasoning capabilities. Moreover, delving deeper into training dynamics, we uncover the roots of saturation and propose a new automatic balancing mechanism to mitigate this limitation. Building on these insights, we propose M-STAR (Multimodal Self-evolving Training for Reasoning), a framework that achieves consistent performance gains across models of varying sizes and diverse benchmarks. All resources are made publicly available at https://mstar-lmm.github.io.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。