arXiv:2503.20211cs.CVcs.RO2025-03CVPR被引 26

用运动与结构先验实现合成到真实场景的鲁棒深度估计

Synthetic-to-Real Self-supervised Robust Depth Estimation via Learning with Motion and Structure Priors

  • 通过代价体中的运动-结构知识迁移,提升合成数据训练效果
  • 在nuScenes和Robotcar上平均提升AbsRel 7.5%、RMSE 4.3%
  • 零样本泛化至雨雾场景,适合户外多天气深度感知任务

单目相机在多种户外条件下(如白天、雨天、夜间)进行自监督深度估计极具挑战,主要源于难以学习通用表征以及真实恶劣数据标签严重缺失。以往方法或依赖合成输入与伪深度标签,或直接套用白天策略,导致效果不佳。本文提出首个合成到真实鲁棒深度估计框架,融合运动与结构先验以有效捕捉真实世界知识。在合成适应阶段,利用冻结的白天模型,在合成恶劣条件下迁移代价体内的运动-结构知识,训练深度估计器;在创新的真实适应阶段,通过一致性重加权策略识别天气无关区域,强化有效伪标签。引入显式深度分布正则化,约束模型在真实数据下的表现。实验表明,本方法在多帧与单帧评估中均优于现有最优水平。在nuScenes与Robotcar数据集上,绝对相对误差(AbsRel)平均提升7.5%,均方根误差(RMSE)平均提升4.3%。在DrivingStereo(雨、雾)的零样本评估中,泛化能力显著优于先前方法。

原文摘要 · Abstract (English)

Self-supervised depth estimation from monocular cameras in diverse outdoor conditions, such as daytime, rain, and nighttime, is challenging due to the difficulty of learning universal representations and the severe lack of labeled real-world adverse data. Previous methods either rely on synthetic inputs and pseudo-depth labels or directly apply daytime strategies to adverse conditions, resulting in suboptimal results. In this paper, we present the first synthetic-to-real robust depth estimation framework, incorporating motion and structure priors to capture real-world knowledge effectively. In the synthetic adaptation, we transfer motion-structure knowledge inside cost volumes for better robust representation, using a frozen daytime model to train a depth estimator in synthetic adverse conditions. In the innovative real adaptation, which targets to fix synthetic-real gaps, models trained earlier identify the weather-insensitive regions with a designed consistency-reweighting strategy to emphasize valid pseudo-labels. We introduce a new regularization by gathering explicit depth distributions to constrain the model when facing real-world data. Experiments show that our method outperforms the state-of-the-art across diverse conditions in multi-frame and single-frame evaluations. We achieve improvements of 7.5% and 4.3% in AbsRel and RMSE on average for nuScenes and Robotcar datasets (daytime, nighttime, rain). In zero-shot evaluation of DrivingStereo (rain, fog), our method generalizes better than the previous ones.

深度估计自监督合成真实迁移多天气泛化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。