用扩散模型提升复杂表面的深度估计,让机器人看得更准
D3RoMa: Disparity Diffusion-based Depth Sensing for Material-Agnostic Robotic Manipulation

- 用去噪扩散模型预测视差图,统一深度估计与修复任务
- 在透明/反光表面场景中实现顶尖精度,真实环境表现优异
- 适合需要高精度深度感知的机器人抓取与操作任务
深度感知是基于3D视觉的机器人技术中的关键问题。然而,现实世界中主动立体或飞行时间(ToF)深度相机常产生噪声大、不完整的问题,严重制约机器人性能。本文提出D3RoMa,一种基于学习的立体图像对深度估计框架,在多样室内场景中可预测清晰准确的深度,尤其在透明或镜面表面等传统方法完全失效的挑战性场景下依然有效。核心思想是将深度估计与修复统一为图像到图像的转换问题,通过去噪扩散概率模型预测视差图,并在推理时引入左右一致性约束作为分类器引导。该框架融合了先进的学习方法与传统立体视觉的几何约束。为训练模型,我们构建了一个包含多种透明和镜面物体的大规模场景级合成数据集,弥补现有桌面数据集的不足。训练后的模型可直接应用于真实世界的自然场景,在多个公开深度估计基准上达到最先进水平。真实环境实验表明,精准深度预测显著提升了机器人在各类场景下的操作性能。
原文摘要 · Abstract (English)
Depth sensing is an important problem for 3D vision-based robotics. Yet, a real-world active stereo or ToF depth camera often produces noisy and incomplete depth which bottlenecks robot performances. In this work, we propose D3RoMa, a learning-based depth estimation framework on stereo image pairs that predicts clean and accurate depth in diverse indoor scenes, even in the most challenging scenarios with translucent or specular surfaces where classical depth sensing completely fails. Key to our method is that we unify depth estimation and restoration into an image-to-image translation problem by predicting the disparity map with a denoising diffusion probabilistic model. At inference time, we further incorporated a left-right consistency constraint as classifier guidance to the diffusion process. Our framework combines recently advanced learning-based approaches and geometric constraints from traditional stereo vision. For model training, we create a large scene-level synthetic dataset with diverse transparent and specular objects to compensate for existing tabletop datasets. The trained model can be directly applied to real-world in-the-wild scenes and achieve state-of-the-art performance in multiple public depth estimation benchmarks. Further experiments in real environments show that accurate depth prediction significantly improves robotic manipulation in various scenarios.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。