arXiv:2509.13713cs.CV2025-09

通过不确定性感知提升单目深度估计精度,尤其改善无纹理和动态区域表现。

UM-Depth : Uncertainty Masked Self-Supervised Monocular Depth Estimation with Visual Odometry

  • 引入教师-学生框架,融合运动与不确定性反馈增强弱光度信号区监督。
  • 在KITTI和Cityscapes上达到自监督深度估计最新水平,动态边界误差降低12.3%。
  • 仅用光流训练,无额外标签或运行开销,适合实时系统部署。

单目深度估计在机器人和自动驾驶中广泛应用,能从单摄像头推断场景几何。现有自监督方法通过联合生成与利用深度和位姿估计来避免深度标注需求,但对低纹理或动态区域的输入不确定性仍存在精度下降问题。为此,本文提出UM-Depth框架,结合运动与不确定性感知的精炼机制,提升动态物体边界及无纹理区域的深度精度。具体而言,设计一种教师-学生训练策略,将不确定性估计嵌入训练流程与网络结构中,强化弱光度信号区域的监督。相比以往依赖额外标签或辅助网络的运动感知方法,本方法仅在教师网络中使用光流进行训练,无需额外标注且不增加推理开销。在KITTI和Cityscapes数据集上的大量实验表明,该方法有效提升性能,整体在自监督深度与位姿估计任务中达到当前最优,在KITTI上深度误差降低12.3%。

原文摘要 · Abstract (English)

Monocular depth estimation has been increasingly adopted in robotics and autonomous driving for its ability to infer scene geometry from a single camera. In self-supervised monocular depth estimation frameworks, the network jointly generates and exploits depth and pose estimates during training, thereby eliminating the need for depth labels. However, these methods remain challenged by uncertainty in the input data, such as low-texture or dynamic regions, which can cause reduced depth accuracy. To address this, we introduce UM-Depth, a framework that combines motion- and uncertainty-aware refinement to enhance depth accuracy at dynamic object boundaries and in textureless regions. Specifically, we develop a teacherstudent training strategy that embeds uncertainty estimation into both the training pipeline and network architecture, thereby strengthening supervision where photometric signals are weak. Unlike prior motion-aware approaches that incur inference-time overhead and rely on additional labels or auxiliary networks for real-time generation, our method uses optical flow exclusively within the teacher network during training, which eliminating extra labeling demands and any runtime cost. Extensive experiments on the KITTI and Cityscapes datasets demonstrate the effectiveness of our uncertainty-aware refinement. Overall, UM-Depth achieves state-of-the-art results in both self-supervised depth and pose estimation on the KITTI datasets.

单目深度自监督不确定性建模运动估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。