CLAP让视频模型跨机器人形态零样本模拟物理世界。
CLAP: Cross-Embodiment Video World Models are Zero-Shot Physical Simulators

- 用末端执行器姿态、语言指令和隐式动作统一不同机器人的控制空间。
- 在DROID等复杂环境中的表现媲美甚至超过单形态顶尖模型。
- 适合想用海量异构视频数据训练通用物理模型的研究者。
当前最先进的动作条件视频模型通常局限于单一机器人形态,无法利用包含丰富物理信号的多样化视频数据。为此,我们提出CLAP框架,可在人类与机器人等多种代理的互联网级视频上训练跨形态动作条件视频生成模型。其核心思想是:无论主体如何,物理规律都支配时空动态。但跨形态学习面临挑战,因不同机器人动作表示差异大,且人类视频中通常无动作标注。CLAP通过三类方式统一动作空间:末端执行器姿态、语言指令与隐式动作。为克服各自局限,提出基于课程学习的跨形态训练策略:先在无标签视频中用隐式动作学习基础物理先验,再将其对齐至末端执行器动作空间,实现零样本部署至真实任务。在DROID等复杂环境中的性能接近或超越现有单形态顶尖模型。该优势可通过少量样本适应进一步放大,开启训练单形态视频世界模型的新范式。最终,CLAP构建了迄今最全面的动作条件视频世界模型集合,涵盖多种动作条件空间(末端执行器、语言、隐式)与机器人形态(跨形态、DROID、Bridge、双臂YAM机器人、G1人形)。代码与模型均已开源。项目官网:https://omni-clap.github.io。
原文摘要 · Abstract (English)
State-of-the-art action-conditioned video models are typically restricted to a single robot embodiment, preventing them from leveraging the vast corpus of heterogeneous video data that contains rich signals for learning generalizable physics. To bridge this gap, we introduce CLAP, a framework for cross-embodiment action-conditioned video generation capable of being trained on diverse, internet-scale videos across human and robotic agents. CLAP is grounded in the insight that universal physical laws govern spatiotemporal dynamics regardless of the actor. However, cross-embodiment learning is non-trivial because action representations vary sharply across robot platforms and are typically absent in human videos. CLAP addresses this fundamental challenge through the following core contributions. First, CLAP reconciles disparate action spaces using end-effector poses, language instructions, and latent actions. Second, to resolve their individual limitations, CLAP introduces a curriculum-based cross-embodiment learning recipe that first learns foundational physical priors across unlabeled video data using latent actions and subsequently grounds them in end-effector action spaces for zero-shot deployment to real-world tasks. Crucially, CLAP approaches or surpasses state-of-the-art single-embodiment video models in challenging environments like DROID. These performance advantages compound via few-shot adaptation to establish a novel paradigm for training single-embodiment video world models. Ultimately, CLAP delivers the most comprehensive suite of action-conditioned video world models to date - spanning diverse action-conditioning spaces (end-effector, language, and latent) and robot morphologies (including cross-embodiment, DROID, Bridge, bimanual YAM robots, and G1 humanoids). We open-source all code and models. Project Website at https://omni-clap.github.io .
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。