用稀疏激光雷达校准深度模型,保持视觉几何一致性。
DrivingDepth: Sparse-Prompted Pixel-wise Scale Correction for Driving Depth Estimation

- 将激光雷达作为几何提示,对冻结的深度模型进行逐像素尺度修正。
- 在nuScenes上实现11.19的AbsRel和5.741的EdgeCR,性能领先。
- 适合需要高精度与几何一致性的自动驾驶深度估计场景。
自动驾驶中的稠密深度估计面临几何-尺度冲突:深度基础模型提供像素对齐的稠密视觉几何,但缺乏可靠度量尺度;而投影的激光雷达虽具度量锚点,却稀疏、噪声大且与图像结构错位。现有稀疏提示方法通过从头重建深度,覆盖了基础模型的连贯几何,导致视觉连续表面上出现结构伪影。我们的关键洞察是:基础模型已捕捉几何连贯的相对深度,无需额外学习表面结构——只需将相对几何映射到度量坐标的一致像素级尺度校正。基于此,我们提出DrivingDepth,将稀疏激光雷达视为几何提示,通过残差像素级尺度校正局部校准冻结的基础先验,从而天然保留稠密视觉几何。在nuScenes上使用四帧环视输入,DrivingDepth达到11.19的AbsRel和5.741的EdgeCR,优于MapAnything(11.99/1.914),同时实现最优度量精度与几何一致性。
原文摘要 · Abstract (English)
Dense depth estimation for autonomous driving faces a geometry-scale conflict: depth foundation models deliver pixel-aligned dense visual geometry without reliable metric scale, while projected LiDAR provides metric anchors that are sparse, noisy, and misaligned with image structures. Existing sparse-prompted methods incorporate LiDAR by regenerating depth from scratch, overriding the foundation model's coherent geometry and producing structural artifacts on visually continuous surfaces. Our key insight is that foundation models already capture geometrically coherent relative depth; no additional surface structure learning is required-only a per-pixel scale factor mapping relative geometry to metric coordinates. Based on this, we propose DrivingDepth, which treats sparse LiDAR as geometric prompts that locally calibrate a frozen foundation prior through residual pixel-wise scale correction, preserving dense visual geometry by construction. On nuScenes with 4-frame surround-view input, DrivingDepth achieves an AbsRel of 11.19 and an EdgeCR of 5.741, outperforming MapAnything (11.99/1.914) by simultaneously delivering SOTA metric accuracy and geometric consistency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。