用规则强化学习提升多模态数学推理能力,性能超越多个主流模型。
MM-Eureka: Exploring the Frontiers of Multimodal Reasoning with Rule-based Reinforcement Learning
- 基于规则的强化学习,在多模态数学数据集上训练7B/32B参数模型。
- 在MMK12数据集上,多模态数学推理准确率显著优于InternVL2.5系列。
- 开源完整代码与数据,适合研究多模态推理与强化学习的开发者。
DeepSeek R1和o1在文本领域通过大规模稳定强化学习展现出强大推理能力。为拓展应用,部分工作尝试将此能力迁移到多模态推理,但受限于任务难度不足与训练规模较小,难以展现强多模态推理能力。为此,我们提出MMK12高质量多模态数学推理数据集,涵盖多样化知识领域,答案与解题过程经人工验证;并构建7B与32B参数的多模态模型MM-EUREKA,采用规则强化学习,结合在线过滤与两阶段训练策略以提升训练稳定性。实验表明,MM-EUREKA在多模态数学推理任务中表现优异,显著超越InternVL2.5-78B及InternVL2.5-38B-MPO等先进模型。其性能在开放源代码模型中具有竞争力,仅略逊于o1在跨学科推理任务中的表现。我们已开源全部代码、模型与数据,详见https://github.com/ModalMinds/MM-EUREKA。
原文摘要 · Abstract (English)
DeepSeek R1, and o1 have demonstrated powerful reasoning capabilities in the text domain through stable large-scale reinforcement learning. To enable broader applications, some works have attempted to transfer these capabilities to multimodal reasoning. However, these efforts have been limited by the limited difficulty of selected tasks and relatively small training scales, making it challenging to demonstrate strong multimodal reasoning abilities. To address this gap, we introduce the MMK12 dataset and MM-EUREKA with 7B and 32B parameters. The former is a high-quality multimodal mathematics reasoning dataset featuring diverse knowledge domains with human-verified answers and solution processes. The latter is a multimodal model employing rule-based reinforcement learning on MMK12, utilizing online filtering and two-stage training strategy to enhance training stability. MM-EUREKA demonstrates remarkable performance gains in multimodal mathematical reasoning, outperforming previous powerful models like InternVL2.5-78B or InternVL2.5-38B-MPO. In particular, MM-EUREKA achieves competitive or superior performance compared to both open-source and closed-source models, and trails slightly behind o1 in multidisciplinary reasoning tasks. We open-source our complete pipeline to foster further research in this area. We release all our codes, models, data, etc. at https://github.com/ModalMinds/MM-EUREKA
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。