让机器人跨形态操作更通用,用共享动态先验+专用控制头实现
DyPES-VLA: Learning Shared Dynamics Priors and Embodiment-Specific Control for Cross-Embodiment Manipulation

- 用未来预测训练视觉语言模型,提取跨机器人的共享动态先验
- 在真实和仿真环境中表现优异,最高达98%成功操作率
- 无需手动对齐动作格式,适配不同机器人本体结构
视觉-语言-动作(VLA)模型已成为机器人操作的强大范式,但为异构机器人本体训练单一通用策略仍是开放问题。现有方法存在两大局限:其一,未能充分利用跨视觉与交互数据的共享动力学先验,限制了跨本体迁移;其二,需大量人工预处理将本体特异性动作转换为统一格式。为此,我们提出 DyPES-VLA,一种学习共享动力学先验与本体特异性控制的跨本体 VLA 模型。首先,在跨本体数据上以未来预测为目标训练视觉语言模型(VLM),使共享查询表示捕捉物体运动、接触及交互引起的场景变化。其次,采用本体特异性混合专家(MoE)动作头,直接将共享动力学先验转化为各本体原生动作空间中的可执行控制,无需手动对齐异构动作。该头共享注意力层以捕捉共性时间动作结构,而本体特异性前馈专家则解决不同本体的运动学约束与控制语义差异。作为通用策略,我们的方法在仿真与真实世界评估中均达到最先进性能,在 LIBERO 上达 98.0% 成功率,RoboCasa-GR1 上达 59.25%,RoboTwin~2.0 上达 89.02%。
原文摘要 · Abstract (English)
Vision-Language-Action (VLA) models have become a powerful paradigm for robot manipulation, but training a single generalist policy for heterogeneous robot embodiments remains an open problem. Existing methods have two main limitations. First, they underuse dynamics priors shared across diverse visual and interaction data, limiting cross-embodiment transfer. Second, they require extensive manual preprocessing to convert embodiment-specific actions into a common format. To overcome these limitations, we propose DyPES-VLA, a cross-embodiment VLA that learns shared Dynamics Priors and Embodiment-Specific control. First, we learn shared dynamics priors by training the vision-language model (VLM) with a future-prediction objective on cross-embodiment data, driving the shared query representation to capture object motion, contact, and interaction-induced scene changes. Second, an embodiment-specific Mixture-of-Experts (MoE) action head translates these shared dynamics priors into executable controls directly in each embodiment's native action space, without manually pre-aligning heterogeneous actions into a common format. This head shares attention layers to capture common temporal action structures, while its embodiment-specific feed-forward experts resolve the unique kinematic constraints and control semantics of distinct embodiments. As a generalist policy, our \ourmethod achieves state-of-the-art performance across simulation and real-world evaluations, reaching 98.0% success on LIBERO, 59.25% on RoboCasa-GR1, and 89.02% on RoboTwin~2.0.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。