arXiv:2602.14974cs.RO2026-02被引 11

DM0让AI从一开始就学会看、说、动,实现物理世界自主决策。

DM0: An Embodied-Native Vision-Language-Action Model towards Physical AI

  • 从零开始统一训练视觉、语言与动作,融合网络文本与真实交互数据。
  • 在RoboChallenge上达到顶尖水平,通用与专用任务表现均领先。
  • 通过空间思维链约束动作空间,适合需要自主操作的机器人研究者。

突破传统互联网预训练模型迁移到物理任务的局限,我们提出DM0,一种面向物理AI的具身原生视觉-语言-动作(VLA)框架。不同于将物理接地视为微调后处理的方式,DM0从一开始便统一学习具身操作与导航。方法采用三阶段流程:预训练、中段训练和后训练。首先,在大规模多源数据上对视觉语言模型(VLM)进行统一预训练,无缝整合网络文本、自动驾驶场景与具身交互日志,联合获取语义知识与物理先验。随后,基于VLM构建流匹配动作专家。为协调高层推理与底层控制,采用混合训练策略:对具身数据,动作专家梯度不反传至VLM以保留泛化表征,而VLM仍可继续学习非具身数据。此外,引入具身空间脚手架策略,构建空间链式思维(CoT)推理,有效约束动作解空间。在RoboChallenge基准测试中,DM0在Table30的专精与通用设置下均取得最先进性能。

原文摘要 · Abstract (English)

Moving beyond the traditional paradigm of adapting internet-pretrained models to physical tasks, we present DM0, an Embodied-Native Vision-Language-Action (VLA) framework designed for Physical AI. Unlike approaches that treat physical grounding as a fine-tuning afterthought, DM0 unifies embodied manipulation and navigation by learning from heterogeneous data sources from the onset. Our methodology follows a comprehensive three-stage pipeline: Pretraining, Mid-Training, and Post-Training. First, we conduct large-scale unified pretraining on the Vision-Language Model (VLM) using diverse corpora--seamlessly integrating web text, autonomous driving scenarios, and embodied interaction logs-to jointly acquire semantic knowledge and physical priors. Subsequently, we build a flow-matching action expert atop the VLM. To reconcile high-level reasoning with low-level control, DM0 employs a hybrid training strategy: for embodied data, gradients from the action expert are not backpropagated to the VLM to preserve generalized representations, while the VLM remains trainable on non-embodied data. Furthermore, we introduce an Embodied Spatial Scaffolding strategy to construct spatial Chain-of-Thought (CoT) reasoning, effectively constraining the action solution space. Experiments on the RoboChallenge benchmark demonstrate that DM0 achieves state-of-the-art performance in both Specialist and Generalist settings on Table30.

具身智能视觉语言动作机器人决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。