用扩散模型生成真实传感器噪声,让仿真机器人直接在真实世界操作
RealD$^2$iff: Bridging Real-World Gap in Robot Manipulation via Depth Diffusion
- 反向思维:从干净深度图生成带真实噪声的图像
- 分层建模全局结构和局部扰动,提升噪声真实性
- 无需真实数据采集,可直接实现零样本仿真到现实迁移
真实世界机器人操作受视觉仿真到现实(sim2real)差距制约,仿真中获取的深度观测无法反映真实传感器的复杂噪声。受扩散模型去噪能力启发,本文提出一种从清洁到嘈杂的新范式,通过学习合成具有真实噪声的深度图来弥合该差距。基于此,我们提出 RealD²iff,一种分层粗到细的扩散框架,将深度噪声分解为全局结构失真与局部细微扰动。为实现对这两类噪声的渐进式建模,我们设计两种互补策略:频域引导监督(FGS)用于全局结构建模,差异引导优化(DGO)用于局部细节优化。构建了涵盖六个阶段的端到端流程,集成于模仿学习框架。实验验证表明,该方法能有效生成类真实深度数据,无需人工收集真实传感器数据;并实现零样本仿真到现实的机器人操控,在不需额外微调的情况下显著提升真实世界表现。
原文摘要 · Abstract (English)
Robot manipulation in the real world is fundamentally constrained by the visual sim2real gap, where depth observations collected in simulation fail to reflect the complex noise patterns inherent to real sensors. In this work, inspired by the denoising capability of diffusion models, we invert the conventional perspective and propose a clean-to-noisy paradigm that learns to synthesize noisy depth, thereby bridging the visual sim2real gap through purely simulation-driven robotic learning. Building on this idea, we introduce RealD$^2$iff, a hierarchical coarse-to-fine diffusion framework that decomposes depth noise into global structural distortions and fine-grained local perturbations. To enable progressive learning of these components, we further develop two complementary strategies: Frequency-Guided Supervision (FGS) for global structure modeling and Discrepancy-Guided Optimization (DGO) for localized refinement. To integrate RealD$^2$iff seamlessly into imitation learning, we construct a pipeline that spans six stages. We provide comprehensive empirical and experimental validation demonstrating the effectiveness of this paradigm. RealD$^2$iff enables two key applications: (1) generating real-world-like depth to construct clean-noisy paired datasets without manual sensor data collection. (2) Achieving zero-shot sim2real robot manipulation, substantially improving real-world performance without additional fine-tuning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。