首次实现稀疏激光雷达注入单目几何模型,显著提升远距离行车深度估计精度。
Sparse-LiDAR Prompting of Monocular Geometry Foundations: An Empirical Study Toward Long-Range Driving Depth

- 设计部分卷积编码器与多尺度融合结构,直接处理真实稀疏激光雷达点云。
- 在100-150米距离上,绝对相对误差降低39%-51%,优于基线模型。
- 支持多种点云密度输入,适合自动驾驶中不规则传感器部署场景。
稀疏激光雷达提示的深度基础模型(PromptDA、Prior Depth Anything、DMD3C)在室内场景或KITTI标准80米评估范围内表现良好,但存在两大局限:(i) 长距离驾驶场景(50-150米)缺乏系统性的分距评估;(ii) 基于视差的基础模型依赖预插值稠密先验,真正稀疏激光雷达注入点图基础模型(如MoGe-2, NeurIPS 2025)尚未探索。本文提出SLIM(Sparse-LiDAR Injected Monocular geometry),首次将MoGe-2适配为接受真实稀疏激光雷达输入。SLIM采用部分卷积稀疏编码器与多尺度融合颈部,在五个尺度上融合激光雷达特征至点图解码器。采用密度无关训练(随机注入率[0.005, 0.30]),使单一模型适应多样输入密度。在Virtual KITTI和CARLA数据集上,SLIM在100-150米距离上将MoGe-2基线的绝对相对误差降低约39%-51%。六种注入率消融实验显示,部分卷积注入在Virtual KITTI上所有设置中均提升AbsRel和RMSE;在CARLA上,五组中AbsRel改善(最接近0.015差异仅0.0013),RMSE在三组中提升(最多0.31单位),其余三组仅下降最多0.11单位。
原文摘要 · Abstract (English)
Sparse-LiDAR-prompted depth foundation models (PromptDA, Prior Depth Anything, DMD3C) have shown strong results on indoor scenes or within KITTI's standard 80-meter evaluation cap. However, two limitations remain: (i) systematic distance-stratified evaluation in long-range driving regimes (50-150 m) is largely absent; (ii) prior approaches built on disparity-based foundations rely on pre-interpolated dense priors, leaving truly sparse LiDAR injection on point-map foundations (e.g., MoGe-2, NeurIPS 2025) unexplored. We present SLIM (Sparse-LiDAR Injected Monocular geometry), the first adaptation of MoGe-2 to accept truly sparse LiDAR input. SLIM integrates a partial-convolution sparse encoder with a multi-scale fusion neck that fuses LiDAR features into the point-map decoder at five scales. We adopt density-agnostic training (random injection ratio in [0.005, 0.30]) so a single model serves diverse input densities. On Virtual KITTI and CARLA, SLIM reduces the absolute relative error of the MoGe-2 baseline by approximately 39-51% at 100-150 m. Ablation across six injection ratios shows partial-convolution injection improves both AbsRel and RMSE on Virtual KITTI in all six settings; on CARLA, AbsRel improves in five of six settings (one near-tie at 0.015 differs by 0.0013), and RMSE is comparable across encoders, with partial-convolution improving in three settings (by up to 0.31 unit) and losing by at most 0.11 unit in the other three.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。