用合成姿势生成视觉数据,提升双臂机器人操作的训练效率
ROPA: Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation
- 用微调的Stable Diffusion生成第三人称视角的RGB-D图像和动作标签
- 2625次仿真与300次真实实验表明性能优于基线方法
- 适合需要大量多样化双臂操作数据的研究者
通过模仿学习训练鲁棒的双臂操作策略,需要覆盖广泛机器人姿态、接触状态和场景上下文的真实演示数据。然而,收集多样且精确的真实演示成本高、耗时长,限制了可扩展性。现有工作虽已采用数据增强,但多针对眼在手上(腕部相机)的RGB输入,或生成无对应动作的新图像,对第三人称视角的RGB-D训练中生成带新动作标签的数据研究较少。本文提出ROP A:面向第三视角RGB-D双臂数据增强的合成机器人姿态生成方法。该方法离线微调Stable Diffusion,生成新型机器人姿态对应的第三人称视角RGB和RGB-D观测,并同时生成对应关节空间动作标签。通过约束优化确保双臂操作中的物理一致性,施加合理的夹持器-物体接触约束。我们在5个模拟任务和3个真实任务上评估该方法,在2625次仿真试验和300次真实试验中验证其优于基线与消融实验,展现出在第三人称双臂操作中可扩展的RGB与RGB-D数据增强潜力。
原文摘要 · Abstract (English)
Training robust bimanual manipulation policies via imitation learning requires demonstration data with broad coverage over robot poses, contacts, and scene contexts. However, collecting diverse and precise real-world demonstrations is costly and time-consuming, which hinders scalability. Prior works have addressed this with data augmentation, typically for either eye-in-hand (wrist camera) setups with RGB inputs or for generating novel images without paired actions, leaving augmentation for eye-to-hand (third-person) RGB-D training with new action labels less explored. In this paper, we propose Synthetic Robot Pose Generation for RGB-D Bimanual Data Augmentation (ROPA), an offline imitation learning data augmentation method that fine-tunes Stable Diffusion to synthesize third-person RGB and RGB-D observations of novel robot poses. Our approach simultaneously generates corresponding joint-space action labels while employing constrained optimization to enforce physical consistency through appropriate gripper-to-object contact constraints in bimanual scenarios. We evaluate our method on 5 simulated and 3 real-world tasks. Our results across 2625 simulation trials and 300 real-world trials demonstrate that ROPA outperforms baselines and ablations, showing its potential for scalable RGB and RGB-D data augmentation in eye-to-hand bimanual manipulation. Our project website is available at: https://ropaaug.github.io/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。