用视觉大模型先验提升单图3D占位预测精度
ViPOcc: Leveraging Visual Priors from Vision Foundation Models for Single-View 3D Occupancy Prediction

- 引入度量深度分支,用逆深度对齐缓解视觉大模型与真实深度的分布差异
- 在KITTI-360和KITTI Raw上实现更准确的3D占位与深度估计
- 适合关注单视角3D场景重建的自动驾驶研究者
从单张图像推断场景三维结构是视觉导向自动驾驶中的病态难题。现有方法通常使用神经辐射场生成体素化3D占位,缺乏实例级语义推理和时序光照一致性。本文提出ViPOcc,利用视觉基础模型(VFMs)的视觉先验实现细粒度3D占位预测。不同于仅依赖体渲染重建RGB与深度图的方法,我们引入度量深度估计分支,提出逆深度对齐模块以弥合VFM预测与真实深度之间的分布差距。恢复的度量深度被用于时序光照对齐与空间几何对齐,确保3D占位预测的准确性与一致性。此外,我们还提出语义引导的非重叠高斯混合采样器,解决先前先进方法中仍存在的冗余与不均衡采样问题。大量实验表明,ViPOcc在KITTI-360和KITTI Raw数据集上的3D占位预测与深度估计任务中均表现优越。
原文摘要 · Abstract (English)
Inferring the 3D structure of a scene from a single image is an ill-posed and challenging problem in the field of vision-centric autonomous driving. Existing methods usually employ neural radiance fields to produce voxelized 3D occupancy, lacking instance-level semantic reasoning and temporal photometric consistency. In this paper, we propose ViPOcc, which leverages the visual priors from vision foundation models (VFMs) for fine-grained 3D occupancy prediction. Unlike previous works that solely employ volume rendering for RGB and depth image reconstruction, we introduce a metric depth estimation branch, in which an inverse depth alignment module is proposed to bridge the domain gap in depth distribution between VFM predictions and the ground truth. The recovered metric depth is then utilized in temporal photometric alignment and spatial geometric alignment to ensure accurate and consistent 3D occupancy prediction. Additionally, we also propose a semantic-guided non-overlapping Gaussian mixture sampler for efficient, instance-aware ray sampling, which addresses the redundant and imbalanced sampling issue that still exists in previous state-of-the-art methods. Extensive experiments demonstrate the superior performance of ViPOcc in both 3D occupancy prediction and depth estimation tasks on the KITTI-360 and KITTI Raw datasets. Our code is available at: \url{https://mias.group/ViPOcc}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。