arXiv:2503.19713cs.ROcs.CV2025-03被引 2

用周边摄像头实现高精度半监督深度估计,解决自动驾驶中尺度模糊问题。

Semi-SMD: Semi-Supervised Metric Depth Estimation via Surrounding Cameras for Autonomous Driving

  • 融合时空语义信息,通过交叉注意力提升深度精度
  • 在DDAD和nuScenes上达到当前最佳性能,误差显著降低
  • 适合自动驾驶多摄像头系统,对标注数据依赖少

本文提出Semi-SMD,一种面向自动驾驶周边摄像头的半监督度量深度估计框架。输入包括相邻帧与相机参数,设计统一的空间-时间-语义融合模块,利用跨摄像头与帧的交叉注意力机制,聚焦度量尺度优化与时序特征匹配。基于周边摄像头、估计深度及外参,构建姿态估计框架,有效缓解多摄像机设置中的尺度模糊问题。同时融合语义世界模型与单目深度模型进行监督,提升深度估计质量。在DDAD与nuScenes数据集上的实验表明,该方法在基于周边摄像头的深度估计任务中达到当前最优性能。源代码将发布于https://github.com/xieyuser/Semi-SMD。

原文摘要 · Abstract (English)

In this paper, we introduce Semi-SMD, a novel metric depth estimation framework tailored for surrounding cameras equipment in autonomous driving. In this work, the input data consists of adjacent surrounding frames and camera parameters. We propose a unified spatial-temporal-semantic fusion module to construct the visual fused features. Cross-attention components for surrounding cameras and adjacent frames are utilized to focus on metric scale information refinement and temporal feature matching. Building on this, we propose a pose estimation framework using surrounding cameras, their corresponding estimated depths, and extrinsic parameters, which effectively address the scale ambiguity in multi-camera setups. Moreover, semantic world model and monocular depth estimation world model are integrated to supervised the depth estimation, which improve the quality of depth estimation. We evaluate our algorithm on DDAD and nuScenes datasets, and the results demonstrate that our method achieves state-of-the-art performance in terms of surrounding camera based depth estimation quality. The source code will be available on https://github.com/xieyuser/Semi-SMD.

深度估计自动驾驶半监督多摄像头

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。