arXiv:2412.06080cs.CVcs.AI2024-12ICCV被引 3

基于垂直位置与物体大小融合,实现车载单目深度零样本准确估计

GVDepth: Zero-Shot Monocular Depth Estimation for Ground Vehicles based on Probabilistic Cue Fusion

  • 构建统一表征解耦相机参数,提升跨数据集泛化能力
  • 在5个自动驾驶数据集上达到与现有零样本方法相当的精度
  • 适合固定摄像头安装的车载场景,无需重新训练

由于单目深度估计固有的病态性,且相机参数与深度之间存在纠缠,导致多数据集训练困难,零样本精度受限。尤其在自动驾驶和移动机器人中,相机固定安装,几何多样性不足。但这一限制也带来机会:相机与地面的固定关系引入额外的透视几何约束,可通过物体在图像中的垂直位置推断深度。然而该线索易过拟合,为此我们提出一种新范式,保持不同相机设置下的一致性,有效解耦深度与具体参数,增强跨数据集泛化能力。同时设计新型架构,自适应地概率融合基于物体大小与垂直位置的深度估计。在五个自动驾驶数据集上的全面评估表明,该方法可实现对多种分辨率、长宽比及相机配置的精确度量深度估计。值得注意的是,仅在一个数据集单相机设置下训练,即达到与现有零样本方法相当的性能。

原文摘要 · Abstract (English)

Generalizing metric monocular depth estimation presents a significant challenge due to its ill-posed nature, while the entanglement between camera parameters and depth amplifies issues further, hindering multi-dataset training and zero-shot accuracy. This challenge is particularly evident in autonomous vehicles and mobile robotics, where data is collected with fixed camera setups, limiting the geometric diversity. Yet, this context also presents an opportunity: the fixed relationship between the camera and the ground plane imposes additional perspective geometry constraints, enabling depth regression via vertical image positions of objects. However, this cue is highly susceptible to overfitting, thus we propose a novel canonical representation that maintains consistency across varied camera setups, effectively disentangling depth from specific parameters and enhancing generalization across datasets. We also propose a novel architecture that adaptively and probabilistically fuses depths estimated via object size and vertical image position cues. A comprehensive evaluation demonstrates the effectiveness of the proposed approach on five autonomous driving datasets, achieving accurate metric depth estimation for varying resolutions, aspect ratios and camera setups. Notably, we achieve comparable accuracy to existing zero-shot methods, despite training on a single dataset with a single-camera setup. Project website: https://unizgfer-lamor.github.io/gvdepth/

单目深度自动驾驶零样本概率融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。