提出地面约束机制,提升单目深度自监督模型的精度与泛化能力。
GroCo: Ground Constraint for Metric Self-Supervised Monocular Depth

- 设计地面区域约束,解决自监督深度估计中的尺度恢复难题。
- 在KITTI上超越现有方法,零样本迁移至DDAD等新数据集表现更优。
- 适合需要高鲁棒性深度估计的自动驾驶场景应用。
近年来单目深度估计取得显著进展,但预测度量深度的模型在不同相机姿态和数据集间仍难以泛化。尽管近期监督方法通过推理时引入地面先验信息缓解此问题,但在自监督设置下因尺度恢复挑战而适应性受限。为此,本文提出一种专为自监督范式设计的地面区域约束机制,不仅能准确恢复尺度,还确保深度预测与地面先验的一致性。实验表明,该方法在KITTI基准上优于现有尺度恢复技术,并显著提升模型泛化能力:在多种相机旋转下表现更稳健,且可在零样本条件下适配此前未见的驾驶数据集(如DDAD)。
原文摘要 · Abstract (English)
Monocular depth estimation has greatly improved in the recent years but models predicting metric depth still struggle to generalize across diverse camera poses and datasets. While recent supervised methods mitigate this issue by leveraging ground prior information at inference, their adaptability to self-supervised settings is limited due to the additional challenge of scale recovery. Addressing this gap, we propose in this paper a novel constraint on ground areas designed specifically for the self-supervised paradigm. This mechanism not only allows to accurately recover the scale but also ensures coherence between the depth prediction and the ground prior. Experimental results show that our method surpasses existing scale recovery techniques on the KITTI benchmark and significantly enhances model generalization capabilities. This improvement can be observed by its more robust performance across diverse camera rotations and its adaptability in zero-shot conditions with previously unseen driving datasets such as DDAD.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。