用动作流统一建模机器人行为,实现跨任务零样本控制。
Hydra-0: Action Flow for Generalist World Modeling and Control

- 将动作转化为像素运动,构建共享视觉接口
- 机器人运动误差降低90.4%,物体运动误差降60.2%
- 支持零样本组合与高效适应,适合多场景机器人控制
我们提出Hydra-0,一种基于动作流的通用世界模型,将机器人动作表示为像素运动。这一共享视觉接口通过学习不同实体、任务、环境及视频生成骨干网络下的动作后果,实现通用世界建模与控制。最佳配置下,机器人运动误差比基准降低90.4%,物体运动误差降低60.2%,同时支持零样本组合与数据高效适应。在RoboLab基准上,重放成功率与参考成功率的相关系数达到r=0.96。此外,我们发现该接口可衍生出逆向模式:从人类示范中转移期望的物体运动流,预测兼容的机器人动作。训练好的动作头将潜在特征映射为可执行动作,无需特定任务的专家机器人演示。这些结果表明,动作流作为共享控制接口,具备连接异构训练数据、开环策略评估与机器人控制的潜力。
原文摘要 · Abstract (English)
We introduce Hydra-0, a generalist world model conditioned on action flow, which represents robot actions as pixel motion. This shared visual interface enables generalist world modeling and control by learning action consequences across embodiments, tasks, environments, and video-generation backbones. Our best configuration achieves 90.4% lower robot-motion error and 60.2% lower object-motion error than our action-conditioned baseline, while supporting zero-shot composition and data-efficient adaptation. On the RoboLab benchmark, Hydra-0 achieves a Pearson correlation of r=0.96 between replayed and reference success rates. Finally, we uncover an emergent inverse mode of this interface: a world action model that predicts compatible robot motion from desired object flow transferred from a human demonstration. A trained action head maps the resulting latent features to executable actions without requiring task-specific expert robot demonstrations. Together, these results demonstrate the potential of action flow as a shared control interface connecting heterogeneous training data, open-loop policy evaluation, and robot control.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。