arXiv:2608.22403cs.RO2026-08

用人类视频训练通用动作模型,让机器人能直接复现动作。

LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models

论文配图:LD4WAM: Learning Latent Dynamics from Human Videos for World Action Models
图 1 · 摘自论文原文
  • 提出运动对齐的潜在动态表征,跨不同机器人形态通用。
  • 在5000小时数据上预训练,生成未来视频并提取动作条件。
  • 适用于抓取与灵巧手机器人,对新物体和背景泛化能力强。

人类视频因其多样性与低成本,在训练世界动作模型(WAM)中日益重要。然而,多数模型仅预测像素级未来帧,无法直接用于动作控制;而运动重定向虽可恢复可执行动作,却在不同机器人形态间存在显著视觉差异。为此,本文提出一种与身体形态无关的运动对齐潜在动态表征,以弥合视频先验与底层动作之间的鸿沟。进一步提出LD4WAM,其包含一个通过语义重建与真实运动对齐训练的潜在动态模型,以及基于混合变换器(MoT)构建的世界动态动作模型。该模型保留完整的未来视频生成能力,并利用可学习查询从生成的未来中提炼潜在动态作为动作条件。在超过5,000小时的人类与机器人数据统一数据集上预训练后,LD4WAM在RoboTwin仿真环境及配备夹持器和灵巧手的真实机器人上表现优异,且对未见过的物体和背景具有良好的泛化能力。

原文摘要 · Abstract (English)

Human video is playing an increasingly central role in training World Action Models (WAMs), owing to its diversity and low collection cost relative to teleoperated robot data. However, most WAMs learn from such video only by predicting pixel-level future frames, giving dynamics that are not directly actionable, whereas motion retargeting recovers directly actionable actions but leaves a large visual gap across embodiments. We therefore propose motion-aligned latent dynamics as an embodiment-agnostic representation to bridge video priors and low-level actions. We further present LD4WAM, which pairs a Latent Dynamics Model trained with semantic reconstruction and real motion alignment with a World Dynamics Action Model built as a mixture-of-transformers (MoT), which preserves full future-video generation and uses learnable queries to distill these latent dynamics from generated futures for action conditioning. Pretrained on our curated unified dataset of over 5{,}000 hours of human and robot data, LD4WAM performs strongly in RoboTwin simulation and on real robots equipped with both grippers and dexterous hands, while generalizing well to unseen objects and backgrounds.

世界动作模型视频生成动作泛化机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。