用仿真数据生成多视角图像,让机器人在真实场景中更稳定操作
Sim2real Image Translation Enables Viewpoint-Robust Policies from Fixed-Camera Datasets
- 基于分割条件的对比损失与强化正则判别器,提升仿真到真实图像转换的一致性
- 仅需少量真实固定视角数据,即可生成未见视角的仿真图像并提升策略性能
- 在桌面上的操作任务中,使视角变化后的成功率提升超40个百分点
基于视觉的机器人操作策略虽取得显著进展,但仍易受相机视角变化等分布偏移影响。真实世界中的机器人示范数据稀缺且视角变化不足。仿真可大规模收集覆盖多视角的示范数据,但存在视觉上的仿真到现实(sim2real)挑战。为此,我们提出MANGO——一种无配对图像翻译方法,包含新颖的分割条件信息归一化互信息损失、高度正则化的判别器设计以及改进的PatchNCE损失。这些组件对保持仿真到真实转换过程中的视角一致性至关重要。训练MANGO时,仅需少量真实世界的固定视角数据,即可通过翻译模拟观测生成多样化的未见视角。在此设定下,MANGO优于所有测试的其他图像翻译方法。在某些真实世界桌面操作任务中,使用MANGO增强后,视角偏移下的成功率达40个百分点以上提升。
原文摘要 · Abstract (English)
Vision-based policies for robot manipulation have achieved significant recent success, but are still brittle to distribution shifts such as camera viewpoint variations. Robot demonstration data is scarce and often lacks appropriate variation in camera viewpoints. Simulation offers a way to collect robot demonstrations at scale with comprehensive coverage of different viewpoints, but presents a visual sim2real challenge. To bridge this gap, we propose MANGO -- an unpaired image translation method with a novel segmentation-conditioned InfoNCE loss, a highly-regularized discriminator design, and a modified PatchNCE loss. We find that these elements are crucial for maintaining viewpoint consistency during sim2real translation. When training MANGO, we only require a small amount of fixed-camera data from the real world, but show that our method can generate diverse unseen viewpoints by translating simulated observations. In this setting, MANGO outperforms all other image translation methods we tested. In certain real-world tabletop manipulation tasks, MANGO augmentation increases shifted-view success rates by over 40 percentage points compared to policies trained without augmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。