用语言模型的思维训练方式,让视觉模型学会深度推理。
Open Vision Reasoner: Transferring Linguistic Cognitive Behavior for Visual Reasoning
- 先用大量语言数据冷启动,再通过近1000步多模态强化学习提升视觉推理能力。
- 在多个基准测试中表现领先,数学推理准确率达95.3%。
- 适合研究多模态推理、行为对齐的学者与开发者。
大型语言模型(LLM)强大的推理能力源于在可验证奖励下的认知行为演化。本文探索如何将这一机制迁移至多模态大模型(MLLM),以激发高级视觉推理能力。基于Qwen2.5-VL-7B,提出两阶段范式:首先进行大规模语言冷启动微调,随后开展近1000步的多模态强化学习(RL),规模超越所有此前开源工作。该研究揭示三大核心发现:1)冷启动阶段即出现行为迁移,源于语言心智意象;2)冷启动广泛记忆视觉行为,而强化学习则关键识别并放大有效模式;3)迁移策略偏好高价值行为,如视觉反思。所提出的Open-Vision-Reasoner(OVR)在多个推理基准上达到顶尖水平,包括MATH500上95.3%、MathVision上51.8%、MathVerse上54.6%。模型、数据与训练动态已公开,以推动更强大且行为对齐的多模态推理器发展。
原文摘要 · Abstract (English)
The remarkable reasoning capability of large language models (LLMs) stems from cognitive behaviors that emerge through reinforcement with verifiable rewards. This work investigates how to transfer this principle to Multimodal LLMs (MLLMs) to unlock advanced visual reasoning. We introduce a two-stage paradigm built on Qwen2.5-VL-7B: a massive linguistic cold-start fine-tuning, followed by multimodal reinforcement learning (RL) spanning nearly 1,000 steps, surpassing all previous open-source efforts in scale. This pioneering work reveals three fundamental insights: 1) Behavior transfer emerges surprisingly early in cold start due to linguistic mental imagery. 2) Cold start broadly memorizes visual behaviors, while RL critically discerns and scales up effective patterns. 3) Transfer strategically favors high-utility behaviors such as visual reflection. Our resulting model, Open-Vision-Reasoner (OVR), achieves state-of-the-art performance on a suite of reasoning benchmarks, including 95.3% on MATH500, 51.8% on MathVision and 54.6% on MathVerse. We release our model, data, and training dynamics to catalyze the development of more capable, behavior-aligned multimodal reasoners.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。