arXiv:2506.03569cs.CL2025-06被引 52

小米开源多模态模型,推理能力超越780亿参数大模型

MiMo-VL Technical Report

  • 四阶段预训练+混合在线强化学习,融合多源奖励信号
  • 在40项任务中35项超越Qwen2.5-VL-7B,奥赛基准达59.4分
  • 适用于GUI识别等场景,性能超专用模型,适合研究与应用

我们开源了MiMo-VL-7B-SFT和MiMo-VL-7B-RL两款强大视觉语言模型,在通用视觉理解与多模态推理上均达到领先水平。MiMo-VL-7B-RL在40项评测任务中有35项优于Qwen2.5-VL-7B,OlympiadBench得分59.4,超越参数量高达780亿的模型。在GUI定位任务中,其在OSWorld-G上取得56.1分,超过专用模型UI-TARS。训练采用四阶段预训练(2.4万亿token)与混合在线强化学习(MORL),整合多样化奖励信号。研究发现,在预训练中引入高质量长链推理数据至关重要,混合强化学习虽面临多领域优化挑战,但收益显著。我们还提供覆盖50余项任务的完整评估套件,以促进可复现性与领域发展。模型检查点与评估套件已开放:https://github.com/XiaomiMiMo/MiMo-VL。

原文摘要 · Abstract (English)

We open-source MiMo-VL-7B-SFT and MiMo-VL-7B-RL, two powerful vision-language models delivering state-of-the-art performance in both general visual understanding and multimodal reasoning. MiMo-VL-7B-RL outperforms Qwen2.5-VL-7B on 35 out of 40 evaluated tasks, and scores 59.4 on OlympiadBench, surpassing models with up to 78B parameters. For GUI grounding applications, it sets a new standard with 56.1 on OSWorld-G, even outperforming specialized models such as UI-TARS. Our training combines four-stage pre-training (2.4 trillion tokens) with Mixed On-policy Reinforcement Learning (MORL) integrating diverse reward signals. We identify the importance of incorporating high-quality reasoning data with long Chain-of-Thought into pre-training stages, and the benefits of mixed RL despite challenges in simultaneous multi-domain optimization. We also contribute a comprehensive evaluation suite covering 50+ tasks to promote reproducibility and advance the field. The model checkpoints and full evaluation suite are available at https://github.com/XiaomiMiMo/MiMo-VL.

多模态视觉语言强化学习模型开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。