arXiv:2601.05839cs.CV2026-01

利用几何一致性提升自动驾驶环视深度估计,性能达到新高度。

GeoSurDepth: Harnessing Foundation Model for Spatial Geometry Consistency-Oriented Self-Supervised Surround-View Depth Estimation

  • 以视觉基础模型为先验,强化三维空间表面法向一致性。
  • 在KITTI、DDAD和nuScenes上实现当前最优效果,精度显著提升。
  • 适合自动驾驶场景下的自监督深度估计研究者参考。

准确的环视深度估计可替代激光传感器,对自动驾驶中的三维场景理解至关重要。现有方法多聚焦于图像级光度约束,较少显式利用单目与环视设置中固有的丰富几何结构。本文提出GeoSurDepth框架,将几何一致性作为核心指导信号。具体而言,利用视觉基础模型作为伪几何先验和特征增强工具,引导网络在三维空间中保持表面法向一致,并在二维空间中实现物体与纹理一致的深度估计。此外,设计了一种新颖的视图合成流程,通过密集深度重建实现2D-3D映射,增强时空上下文下的光度监督,弥补目标视图图像重建的不足。最后,提出自适应联合运动学习策略,使网络能动态强化有效的空间几何线索,提升运动推理能力。在KITTI、DDAD和nuScenes上的大量实验表明,GeoSurDepth达到当前最优性能,验证了该方法的有效性。本框架凸显了挖掘几何相干性与一致性的关键作用,为鲁棒的自监督深度估计提供了新思路。

原文摘要 · Abstract (English)

Accurate surround-view depth estimation provides a competitive alternative to laser-based sensors and is essential for 3D scene understanding in autonomous driving. While empirical studies have proposed various approaches that primarily focus on enforcing cross-view constraints at photometric level, few explicitly exploit the rich geometric structure inherent in both monocular and surround-view setting. In this work, we propose GeoSurDepth, a framework that leverages geometry consistency as the primary cue for surround-view depth estimation. Concretely, we utilize vision foundation models as pseudo geometry priors and feature representation enhancement tool to guide the network to maintain surface normal consistency in spatial 3D space and regularize object- and texture-consistent depth estimation in 2D. In addition, we introduce a novel view synthesis pipeline where 2D-3D lifting is achieved with dense depth reconstructed via spatial warping, encouraging additional photometric supervision across temporal and spatial contexts, and compensating for the limitations of target-view image reconstruction. Finally, a newly-proposed adaptive joint motion learning strategy enables the network to adaptively emphasize informative spatial geometry cues for improved motion reasoning. Extensive experiments on KITTI, DDAD and nuScenes demonstrate that GeoSurDepth achieves SoTA performance, validating the effectiveness of our approach. Our framework highlights the importance of exploiting geometry coherence and consistency for robust self-supervised depth estimation.

深度估计几何一致性自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。