用激光雷达引导多视角立体视觉,提升自动驾驶深度估计精度与稳定性
LiDAR Prompted Spatio-Temporal Multi-View Stereo for Autonomous Driving
- 以激光雷达点云作为几何先验,锚定深度绝对尺度
- 融合多线索并用时空解码器,实现帧间深度一致性
- 在跨域零样本迁移中表现优异,适合实际自动驾驶系统
准确的度量深度对自动驾驶感知与仿真至关重要,但现有方法难以同时实现高精度、多视角与时间一致性以及跨域泛化。为此,我们提出DriveMVS,一种新型多视角立体框架,通过两个核心洞察解决这些矛盾目标:(1) 稀疏但度量准确的激光雷达观测可作为几何提示,锚定深度估计的绝对尺度;(2) 深度融合多样线索有助于解决歧义并增强鲁棒性,而时空解码器确保帧间一致性。基于此,DriveMVS以两种方式嵌入激光雷达提示:作为硬几何先验锚定代价体,以及通过三线索组合器进行软特征级融合。在时间一致性方面,采用时空解码器联合利用多视角代价体的几何线索与邻近帧的时间上下文。实验表明,DriveMVS在多个基准上达到顶尖性能,在度量精度、时间稳定性及零样本跨域迁移方面表现突出,展示了其在可扩展、可靠自动驾驶系统中的实用价值。
原文摘要 · Abstract (English)
Accurate metric depth is critical for autonomous driving perception and simulation, yet current approaches struggle to achieve high metric accuracy, multi-view and temporal consistency, and cross-domain generalization. To address these challenges, we present DriveMVS, a novel multi-view stereo framework that reconciles these competing objectives through two key insights: (1) Sparse but metrically accurate LiDAR observations can serve as geometric prompts to anchor depth estimation in absolute scale, and (2) deep fusion of diverse cues is essential for resolving ambiguities and enhancing robustness, while a spatio-temporal decoder ensures consistency across frames. Built upon these principles, DriveMVS embeds the LiDAR prompt in two ways: as a hard geometric prior that anchors the cost volume, and as soft feature-wise guidance fused by a triple-cue combiner. Regarding temporal consistency, DriveMVS employs a spatio-temporal decoder that jointly leverages geometric cues from the MVS cost volume and temporal context from neighboring frames. Experiments show that DriveMVS achieves state-of-the-art performance on multiple benchmarks, excelling in metric accuracy, temporal stability, and zero-shot cross-domain transfer, demonstrating its practical value for scalable, reliable autonomous driving systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。