用视频扩散模型生成多样化的双臂机器人操作数据,提升学习泛化能力。
CRAFT: Video Diffusion for Bimanual Robot Data Generation
- 基于边缘结构引导的视频扩散框架,生成时序一致的操作视频。
- 仅需少量真实示范,即可合成大量视觉多样且物理合理的训练数据。
- 支持视角、光照、物体姿态等多维度增强,适合双臂机器人学习场景。
从示范中学习双臂机器人操作,受限于真实数据成本高、视觉多样性差,导致策略在不同视角、物体配置和机器人形态下鲁棒性不足。我们提出基于视频扩散变换器的边缘引导机器人数据生成方法(CRAFT),通过模拟轨迹提取的边缘结构作为条件,生成时序连贯的操作视频并附带动作标签。该方法可统一实现物体位姿变化、相机视角、光照与背景差异、跨形态迁移及多视角合成等增强。利用预训练视频扩散模型,将模拟视频与仿真轨迹中的动作标签转换为动作一致的示范数据。仅需少数真实示范,即可生成大规模、高视觉多样性的逼真训练数据,无需在真实机器人上重播演示(Sim2Real)。在模拟与真实双臂任务中,相比现有增强策略与单纯数据扩展,CRAFT显著提升成功率,证明扩散模型可有效拓展示范多样性,改善双臂操作任务的泛化性能。
原文摘要 · Abstract (English)
Bimanual robot learning from demonstrations is fundamentally limited by the cost and narrow visual diversity of real-world data, which constrains policy robustness across viewpoints, object configurations, and embodiments. We present Canny-guided Robot Data Generation using Video Diffusion Transformers (CRAFT), a video diffusion-based framework for scalable bimanual demonstration generation that synthesizes temporally coherent manipulation videos while producing action labels. By conditioning video diffusion on edge-based structural cues extracted from simulator-generated trajectories, CRAFT produces physically plausible trajectory variations and supports a unified augmentation pipeline spanning object pose changes, camera viewpoints, lighting and background variations, cross-embodiment transfer, and multi-view synthesis. We leverage a pre-trained video diffusion model to convert simulated videos, along with action labels from the simulation trajectories, into action-consistent demonstrations. Starting from only a few real-world demonstrations, CRAFT generates a large, visually diverse set of photorealistic training data, bypassing the need to replay demonstrations on the real robot (Sim2Real). Across simulated and real-world bimanual tasks, CRAFT improves success rates over existing augmentation strategies and straightforward data scaling, demonstrating that diffusion-based video generation can substantially expand demonstration diversity and improve generalization for dual-arm manipulation tasks. Our project website is available at: https://craftaug.github.io/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。