arXiv:2604.02639cs.CVcs.AI2026-04

针对铰接式车辆设计自监督环视深度估计新方法,提升复杂结构下的深度精度。

Cross-Vehicle 3D Geometric Consistency for Self-Supervised Surround Depth Estimation on Articulated Vehicles

  • 利用多视角上下文与表面法向约束增强空间时序一致性
  • 通过地面感知与相机高度正则化实现度量级深度估计
  • 适用于自动驾驶铰接车辆,适配性强且性能领先

环视深度估计为自动驾驶提供了一种低成本的3D感知替代方案。尽管近期自监督方法在多相机设置下提升了尺度感知与场景覆盖能力,但主要针对乘用车,极少考虑铰接式车辆或机器人平台。其复杂跨段几何结构与运动耦合使跨视角深度推理更具挑战。本文提出 extbf{ArticuSurDepth},一种面向铰接式车辆的自监督环视深度估计框架,通过视觉基础模型提供的结构先验,引导跨视角与跨车辆几何一致性来增强深度学习。具体包括多视角空间上下文增强策略和跨视角表面法向约束,以提升时空上下文中的结构一致性;进一步引入相机高度正则化与地面感知机制,促进度量深度估计,并通过跨车辆位姿一致性连接各铰接段的运动估计。为验证方法,构建了铰接式车辆实验平台并采集数据集。实验结果表明,在自建数据集以及DDAD、nuScenes、KITTI基准上均达到当前最优(SoTA)性能。

原文摘要 · Abstract (English)

Surround depth estimation provides a cost-effective alternative to LiDAR for 3D perception in autonomous driving. While recent self-supervised methods explore multi-camera settings to improve scale awareness and scene coverage, they are primarily designed for passenger vehicles and rarely consider articulated vehicles or robotics platforms. The articulated structure introduces complex cross-segment geometry and motion coupling, making consistent depth reasoning across views more challenging. In this work, we propose \textbf{ArticuSurDepth}, a self-supervised framework for surround-view depth estimation on articulated vehicles that enhances depth learning through cross-view and cross-vehicle geometric consistency guided by structural priors from vision foundation model. Specifically, we introduce multi-view spatial context enrichment strategy and a cross-view surface normal constraint to improve structural coherence across spatial and temporal contexts. We further incorporate camera height regularization with ground plane-awareness to encourage metric depth estimation, together with cross-vehicle pose consistency that bridges motion estimation between articulated segments. To validate our proposed method, an articulated vehicle experiment platform was established with a dataset collected over it. Experiment results demonstrate state-of-the-art (SoTA) performance of depth estimation on our self-collected dataset as well as on DDAD, nuScenes, and KITTI benchmarks.

深度估计自监督自动驾驶铰接车辆

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。