用扩散模型提升非朗伯物体深度补全,增强机器人感知与抓取能力
DidSee: Diffusion-Based Depth Completion for Material-Agnostic Robotic Perception and Manipulation
- 引入零终态信噪比噪声调度,消除信号泄漏偏差
- 单步训练减少误差累积,结合任务专用损失优化性能
- 融合语义增强模块,实现深度补全与分割联合优化
商用RGB-D相机对非朗伯物体常生成噪声大、不完整的深度图。传统深度补全方法因训练数据多样性与规模有限而泛化能力差。近期工作利用预训练文本到图像扩散模型的视觉先验提升密集预测泛化性,但我们发现原始扩散框架中训练-推理不匹配导致显著偏差,且非朗伯区域视觉特征模糊进一步制约精确预测。为此,提出 extbf{DidSee}:首先,采用重缩放噪声调度器,强制终端信噪比为零,消除信号泄漏;其次,设计无噪声依赖的单步训练方式,缓解暴露偏差并用任务特定损失优化模型;最后,引入语义增强模块,实现深度补全与语义分割联合优化,区分物体与背景,生成精细深度图。在多个基准测试中达领先性能,展现强现实泛化能力,并有效提升类别级位姿估计与机器人抓取等下游任务效果。
原文摘要 · Abstract (English)
Commercial RGB-D cameras often produce noisy, incomplete depth maps for non-Lambertian objects. Traditional depth completion methods struggle to generalize due to the limited diversity and scale of training data. Recent advances exploit visual priors from pre-trained text-to-image diffusion models to enhance generalization in dense prediction tasks. However, we find that biases arising from training-inference mismatches in the vanilla diffusion framework significantly impair depth completion performance. Additionally, the lack of distinct visual features in non-Lambertian regions further hinders precise prediction. To address these issues, we propose \textbf{DidSee}, a diffusion-based framework for depth completion on non-Lambertian objects. First, we integrate a rescaled noise scheduler enforcing a zero terminal signal-to-noise ratio to eliminate signal leakage bias. Second, we devise a noise-agnostic single-step training formulation to alleviate error accumulation caused by exposure bias and optimize the model with a task-specific loss. Finally, we incorporate a semantic enhancer that enables joint depth completion and semantic segmentation, distinguishing objects from backgrounds and yielding precise, fine-grained depth maps. DidSee achieves state-of-the-art performance on multiple benchmarks, demonstrates robust real-world generalization, and effectively improves downstream tasks such as category-level pose estimation and robotic grasping.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。