arXiv:2512.11797cs.ROcs.CV2025-12被引 9

用视频扩散模型生成符合机器人动作的仿真数据,解决真实数据难获取问题。

AnchorDream: Repurposing Video Diffusion for Embodiment-Aware Robot Data Synthesis

  • 以机器人运动图为条件,控制扩散过程生成一致动作
  • 仅需少量人类操作示范,生成大规模高质量数据集
  • 在模拟与真实场景中分别提升36.4%和近一倍性能

大规模多样化的机器人示范数据仍是模仿学习的主要瓶颈,因真实世界数据采集成本高,而模拟器存在多样性与保真度不足、明显的仿真到现实差距。生成模型虽具吸引力,但现有方法通常仅改变视觉外观,未生成新行为,或因身体一致性问题导致不合理动作。为此,我们提出AnchorDream,一种基于机器人动作的感知世界模型,复用预训练视频扩散模型进行机器人数据合成。该模型以机器人运动渲染图为条件,锚定本体特征,防止幻觉,同时生成与机器人运动学一致的物体与环境。仅需少量人类远程操控示范,即可扩展为大规模、多样化、高质量数据集,无需显式环境建模。实验表明,生成数据在下游策略学习中表现优异:模拟器基准提升36.4%相对收益,真实场景中性能接近翻倍。结果表明,将生成世界模型与机器人动作结合,是规模化模仿学习的可行路径。

原文摘要 · Abstract (English)

The collection of large-scale and diverse robot demonstrations remains a major bottleneck for imitation learning, as real-world data acquisition is costly and simulators offer limited diversity and fidelity with pronounced sim-to-real gaps. While generative models present an attractive solution, existing methods often alter only visual appearances without creating new behaviors, or suffer from embodiment inconsistencies that yield implausible motions. To address these limitations, we introduce AnchorDream, an embodiment-aware world model that repurposes pretrained video diffusion models for robot data synthesis. AnchorDream conditions the diffusion process on robot motion renderings, anchoring the embodiment to prevent hallucination while synthesizing objects and environments consistent with the robot's kinematics. Starting from only a handful of human teleoperation demonstrations, our method scales them into large, diverse, high-quality datasets without requiring explicit environment modeling. Experiments show that the generated data leads to consistent improvements in downstream policy learning, with relative gains of 36.4% in simulator benchmarks and nearly double performance in real-world studies. These results suggest that grounding generative world models in robot motion provides a practical path toward scaling imitation learning.

机器人数据合成视频扩散模型模仿学习仿真到现实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。