arXiv:2507.14845cs.CV2025-07

仅用稀疏深度和单张图像实现自监督深度补全

Training Self-Supervised Depth Completion Using Sparse Measurements and a Single Image

  • 基于深度分布特性设计新损失函数,从观测点传播深度信息
  • 无需密集标注或多帧图像,支持静态场景训练
  • 融合视觉基础模型分割图,提升深度估计精度

深度补全是重要视觉任务,旨在从稀疏深度测量中恢复稠密深度图。现有方法中,监督学习依赖密集深度标签,自监督方法需多帧图像以保证几何约束和光度一致性,但密集标注成本高,多帧依赖限制了其在静态或单帧场景的应用。为此,本文提出一种新型自监督深度补全范式,仅需稀疏深度测量及其对应图像即可训练。不同于以往方法,该方法无需密集标签或邻近视角图像。通过利用深度分布特性,设计新型损失函数,有效将深度信息从已观测点传播至未观测区域。同时,引入视觉基础模型生成的分割图以进一步提升深度估计性能。大量实验验证了所提方法的有效性。

原文摘要 · Abstract (English)

Depth completion is an important vision task, and many efforts have been made to enhance the quality of depth maps from sparse depth measurements. Despite significant advances, training these models to recover dense depth from sparse measurements remains a challenging problem. Supervised learning methods rely on dense depth labels to predict unobserved regions, while self-supervised approaches require image sequences to enforce geometric constraints and photometric consistency between frames. However, acquiring dense annotations is costly, and multi-frame dependencies limit the applicability of self-supervised methods in static or single-frame scenarios. To address these challenges, we propose a novel self-supervised depth completion paradigm that requires only sparse depth measurements and their corresponding image for training. Unlike existing methods, our approach eliminates the need for dense depth labels or additional images captured from neighboring viewpoints. By leveraging the characteristics of depth distribution, we design novel loss functions that effectively propagate depth information from observed points to unobserved regions. Additionally, we incorporate segmentation maps generated by vision foundation models to further enhance depth estimation. Extensive experiments demonstrate the effectiveness of our proposed method.

深度补全自监督单图像

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。