通过分阶段强化学习,让小模型学会跨模态推理。
Infi-MMR: Curriculum-based Unlocking Multimodal Reasoning via Phased Reinforcement Learning in Multimodal Small Language Models
- 三阶段课程式训练:从文本到多模态,逐步激活推理能力。
- 在多个数学推理数据集上达到领先表现,最高达43.68%准确率。
- 适合研究小模型多模态推理或需要轻量化推理系统的人参考。
近期大语言模型(LLMs)在推理能力方面取得显著进展,如DeepSeek-R1通过基于规则的强化学习显著提升逻辑推理能力。然而,将这些成果扩展至多模态大语言模型(MLLMs)面临关键挑战,尤其对多模态小语言模型(MSLMs)而言,其基础推理能力较弱:(1)高质量多模态推理数据稀缺;(2)视觉模块集成导致推理能力下降;(3)直接应用强化学习可能生成复杂但错误的推理过程。为此,我们提出新框架Infi-MMR,通过三个精心设计的阶段系统性释放MSLMs的推理潜力,并构建了多模态推理模型Infi-MMR-3B。第一阶段“基础推理激活”利用高质量文本推理数据强化逻辑能力;第二阶段“跨模态推理适配”使用带字幕的多模态数据促进推理技能向多模态场景迁移;第三阶段“多模态推理增强”采用无字幕的精选多模态数据,缓解语言偏见,提升鲁棒性跨模态推理。Infi-MMR-3B在多个任务上达到当前最优:MathVerse testmini为43.68%,MathVision test为27.04%,OlympiadBench为21.33%,且在MathVista testmini上通用推理能力达67.2%。资源已公开于https://huggingface.co/Reallm-Labs/Infi-MMR-3B。
原文摘要 · Abstract (English)
Recent advancements in large language models (LLMs) have demonstrated substantial progress in reasoning capabilities, such as DeepSeek-R1, which leverages rule-based reinforcement learning to enhance logical reasoning significantly. However, extending these achievements to multimodal large language models (MLLMs) presents critical challenges, which are frequently more pronounced for Multimodal Small Language Models (MSLMs) given their typically weaker foundational reasoning abilities: (1) the scarcity of high-quality multimodal reasoning datasets, (2) the degradation of reasoning capabilities due to the integration of visual processing, and (3) the risk that direct application of reinforcement learning may produce complex yet incorrect reasoning processes. To address these challenges, we design a novel framework Infi-MMR to systematically unlock the reasoning potential of MSLMs through a curriculum of three carefully structured phases and propose our multimodal reasoning model Infi-MMR-3B. The first phase, Foundational Reasoning Activation, leverages high-quality textual reasoning datasets to activate and strengthen the model's logical reasoning capabilities. The second phase, Cross-Modal Reasoning Adaptation, utilizes caption-augmented multimodal data to facilitate the progressive transfer of reasoning skills to multimodal contexts. The third phase, Multimodal Reasoning Enhancement, employs curated, caption-free multimodal data to mitigate linguistic biases and promote robust cross-modal reasoning. Infi-MMR-3B achieves both state-of-the-art multimodal math reasoning ability (43.68% on MathVerse testmini, 27.04% on MathVision test, and 21.33% on OlympiadBench) and general reasoning ability (67.2% on MathVista testmini). Resources are available at https://huggingface.co/Reallm-Labs/Infi-MMR-3B.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。