用扩散模型生成双臂协同的视觉与动作数据,提升机器人模仿学习效率。
D-CODA: Diffusion for Coordinated Dual-Arm Data Augmentation
- 基于扩散模型同步生成双腕摄像头视角图像和对应关节动作。
- 在2250次仿真与300次真实实验中均优于基线方法。
- 适合需要高效采集双臂协同数据的研究者或工业应用。
双臂操作学习因高维状态空间和双臂间紧密协调而具有挑战性。眼手模仿学习通过腕部摄像头聚焦任务相关视图简化感知,但多样示范数据收集成本高昂,亟需可扩展的数据增强方法。现有视觉增强多集中于单臂场景,而双臂操作需同时生成两臂视角一致的观测结果,并生成有效且可行的动作标签。本文提出面向眼手双臂模仿学习的离线数据增强方法 D-CODA,训练扩散模型合成双臂腕部相机图像,同时生成关节空间动作标签。采用约束优化确保涉及夹爪-物体接触的状态满足双臂协调约束。我们在5个模拟任务和3个真实任务上评估D-CODA,2250次仿真试验与300次真实试验结果表明其性能超越基线及消融实验,展现了在眼手双臂操作中可扩展数据增强的潜力。
原文摘要 · Abstract (English)
Learning bimanual manipulation is challenging due to its high dimensionality and tight coordination required between two arms. Eye-in-hand imitation learning, which uses wrist-mounted cameras, simplifies perception by focusing on task-relevant views. However, collecting diverse demonstrations remains costly, motivating the need for scalable data augmentation. While prior work has explored visual augmentation in single-arm settings, extending these approaches to bimanual manipulation requires generating viewpoint-consistent observations across both arms and producing corresponding action labels that are both valid and feasible. In this work, we propose Diffusion for COordinated Dual-arm Data Augmentation (D-CODA), a method for offline data augmentation tailored to eye-in-hand bimanual imitation learning that trains a diffusion model to synthesize novel, viewpoint-consistent wrist-camera images for both arms while simultaneously generating joint-space action labels. It employs constrained optimization to ensure that augmented states involving gripper-to-object contacts adhere to constraints suitable for bimanual coordination. We evaluate D-CODA on 5 simulated and 3 real-world tasks. Our results across 2250 simulation trials and 300 real-world trials demonstrate that it outperforms baselines and ablations, showing its potential for scalable data augmentation in eye-in-hand bimanual manipulation. Our project website is at: https://dcodaaug.github.io/D-CODA/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。