用仿真数据训练的深度模型,让机器人在真实世界精准抓取复杂物体。
Manipulation as in Simulation: Enabling Accurate Geometry Perception in Robots
- 通过仿真生成带噪声的深度数据,训练出能去噪的深度模型。
- 在真实机器人上完成长序列任务,性能几乎无损失。
- 无需额外调优,直接从仿真迁移到真实世界,适合通用机器人研究。
现代机器人操作主要依赖2D彩色图像进行技能学习,但泛化能力差。人类则更依赖距离、尺寸和形状等3D几何属性。尽管深度相机可获取此类信息,但其精度低且易受噪声影响。本文提出相机深度模型(CDMs),作为日常深度相机的轻量级插件,输入RGB图像与原始深度信号,输出去噪后的精确度量深度。为此,我们构建了神经数据引擎,通过模拟深度相机的噪声模式生成高质量配对数据。实验表明,CDMs在深度预测上接近仿真级别精度,有效弥合了仿真到现实的差距。特别地,首次证明:仅用原始仿真深度训练的策略,无需加噪声或现实微调,即可无缝应用于两个涉及可动、反光、细长物体的长时序真实任务,性能几乎无下降。希望推动未来利用仿真数据与3D信息构建通用机器人策略的研究。
原文摘要 · Abstract (English)
Modern robotic manipulation primarily relies on visual observations in a 2D color space for skill learning but suffers from poor generalization. In contrast, humans, living in a 3D world, depend more on physical properties-such as distance, size, and shape-than on texture when interacting with objects. Since such 3D geometric information can be acquired from widely available depth cameras, it appears feasible to endow robots with similar perceptual capabilities. Our pilot study found that using depth cameras for manipulation is challenging, primarily due to their limited accuracy and susceptibility to various types of noise. In this work, we propose Camera Depth Models (CDMs) as a simple plugin on daily-use depth cameras, which take RGB images and raw depth signals as input and output denoised, accurate metric depth. To achieve this, we develop a neural data engine that generates high-quality paired data from simulation by modeling a depth camera's noise pattern. Our results show that CDMs achieve nearly simulation-level accuracy in depth prediction, effectively bridging the sim-to-real gap for manipulation tasks. Notably, our experiments demonstrate, for the first time, that a policy trained on raw simulated depth, without the need for adding noise or real-world fine-tuning, generalizes seamlessly to real-world robots on two challenging long-horizon tasks involving articulated, reflective, and slender objects, with little to no performance degradation. We hope our findings will inspire future research in utilizing simulation data and 3D information in general robot policies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。