首个跨具身基础模型,同时突破自动驾驶与具身智能瓶颈
MiMo-Embodied: X-Embodied Foundation Model Technical Report
- 多阶段学习+精心构建数据集,融合驾驶与具身任务
- 17个具身智能基准、12个自动驾驶基准全破纪录
- 适合研究跨领域通用智能的学者与工程师
我们开源了MiMo-Embodied,首个在自主驾驶和具身智能两个领域均实现顶尖性能的跨具身基础模型。该模型在17个具身智能基准(涵盖任务规划、可及性预测、空间理解)上创下新纪录,同时在12个自动驾驶基准(包括环境感知、状态预测、驾驶规划)中表现卓越。通过多阶段训练、精心设计的数据构建以及思维链/强化学习微调,两个领域展现出显著正向迁移和相互增强效应。本报告详细分析了模型架构与训练方法,推动后续研究发展。代码与模型已公开于 https://github.com/XiaomiMiMo/MiMo-Embodied。
原文摘要 · Abstract (English)
We open-source MiMo-Embodied, the first cross-embodied foundation model to successfully integrate and achieve state-of-the-art performance in both Autonomous Driving and Embodied AI. MiMo-Embodied sets new records across 17 embodied AI benchmarks in Task Planning, Affordance Prediction and Spatial Understanding, while also excelling in 12 autonomous driving benchmarks across Environmental Perception, Status Prediction, and Driving Planning. Across these tasks, MiMo-Embodied significantly outperforms existing open-source, closed-source, and specialized baselines. Our results indicate that through multi-stage learning, curated data construction, and CoT/RL fine-tuning, these two domains exhibit strong positive transfer and mutually reinforce one another. We provide a detailed analysis of our model design and training methodologies to facilitate further research. Code and models are available at https://github.com/XiaomiMiMo/MiMo-Embodied.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。