arXiv:2505.12705cs.ROcs.AI2025-05被引 146

用视频模型生成机器人数据,仅靠一次操作就能学会22种新动作。

DreamGen: Unlocking Generalization in Robot Learning through Video World Models

  • 通过视频生成模型合成机器人动作数据,再还原伪动作序列。
  • 仅需一个场景的抓取数据,即可在已知和未知环境中完成22种新行为。
  • 适合希望减少实操数据收集的机器人学习研究者。

我们提出DreamGen,一种四阶段训练机器人策略的简单但高效的方法,通过神经轨迹——即由视频世界模型生成的合成机器人数据——实现跨行为与环境的泛化。DreamGen利用先进的图像到视频生成模型,适配目标机器人形态,生成熟悉或新任务在多样化环境中的逼真合成视频。由于这些模型仅生成视频,我们通过潜在动作模型或逆动力学模型(IDM)恢复伪动作序列。尽管方法简单,DreamGen实现了强大的行为与环境泛化:一个人形机器人可在已见和未见环境中执行22种新行为,且仅需在单一环境中获取一次抓取-放置任务的遥操作数据。为系统评估该流程,我们引入DreamGen Bench,一个视频生成基准,结果显示其性能与下游策略成功率高度相关。本工作开辟了机器人学习规模化的新路径,远超人工数据采集的局限。代码已开源于https://github.com/NVIDIA/GR00T-Dreams。

原文摘要 · Abstract (English)

We introduce DreamGen, a simple yet highly effective 4-stage pipeline for training robot policies that generalize across behaviors and environments through neural trajectories - synthetic robot data generated from video world models. DreamGen leverages state-of-the-art image-to-video generative models, adapting them to the target robot embodiment to produce photorealistic synthetic videos of familiar or novel tasks in diverse environments. Since these models generate only videos, we recover pseudo-action sequences using either a latent action model or an inverse-dynamics model (IDM). Despite its simplicity, DreamGen unlocks strong behavior and environment generalization: a humanoid robot can perform 22 new behaviors in both seen and unseen environments, while requiring teleoperation data from only a single pick-and-place task in one environment. To evaluate the pipeline systematically, we introduce DreamGen Bench, a video generation benchmark that shows a strong correlation between benchmark performance and downstream policy success. Our work establishes a promising new axis for scaling robot learning well beyond manual data collection. Code available at https://github.com/NVIDIA/GR00T-Dreams.

机器人学习视频生成泛化能力合成数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。