统一视觉、动作与世界建模,提升机器人任务表现
Motus: A Unified Latent Action World Model

- 采用MoT架构融合理解、视频生成与动作专家,支持多模式切换
- 在仿真中比X-VLA高15%,比Pi0.5高45%,真实场景提升11%~48%
- 利用光流学习潜空间动作,支持大规模动作预训练
当前具身智能体通常依赖孤立的模型完成感知、世界建模与控制,导致难以整合多模态生成能力并限制从异构数据中学习。本文提出Motus,一种统一的潜在动作世界模型,利用现有预训练通用模型和丰富的可共享运动信息。Motus引入混合变压器(MoT)架构,集成理解、视频生成与动作三个专家,并采用类似UniDiffuser的调度器,实现世界模型、视觉-语言-动作模型、逆动力学模型、视频生成模型及视频-动作联合预测模型间的灵活切换。进一步地,通过光流学习潜在动作,设计三阶段训练流程与六层数据金字塔,提取像素级“动作差值”,支持大规模动作预训练。实验表明,Motus在仿真中相较SOTA方法提升15%(优于X-VLA)和45%(优于Pi0.5),在真实场景中提升11%~48%,证明统一建模所有功能与先验显著提升下游机器人任务性能。
原文摘要 · Abstract (English)
While a general embodied agent must function as a unified system, current methods are built on isolated models for understanding, world modeling, and control. This fragmentation prevents unifying multimodal generative capabilities and hinders learning from large-scale, heterogeneous data. In this paper, we propose Motus, a unified latent action world model that leverages existing general pretrained models and rich, sharable motion information. Motus introduces a Mixture-of-Transformer (MoT) architecture to integrate three experts (i.e., understanding, video generation, and action) and adopts a UniDiffuser-style scheduler to enable flexible switching between different modeling modes (i.e., world models, vision-language-action models, inverse dynamics models, video generation models, and video-action joint prediction models). Motus further leverages the optical flow to learn latent actions and adopts a recipe with three-phase training pipeline and six-layer data pyramid, thereby extracting pixel-level "delta action" and enabling large-scale action pretraining. Experiments show that Motus achieves superior performance against state-of-the-art methods in both simulation (a +15% improvement over X-VLA and a +45% improvement over Pi0.5) and real-world scenarios(improved by +11~48%), demonstrating unified modeling of all functionalities and priors significantly benefits downstream robotic tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。