arXiv:2409.13430cs.CVcs.AI2024-09ECCV被引 29

利用时间几何对应关系提升单目3D占位预测精度

CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction

论文配图:CVT-Occ: Cost Volume Temporal Fusion for 3D Occupancy Prediction
图 1 · 摘自论文原文
  • 通过历史帧点特征构建时序代价体,融合多帧信息
  • 在Occ3D-Waymo数据集上超越当前最优方法
  • 计算开销小,适合自动驾驶实时系统

基于视觉的3D占位预测受单目视觉深度估计固有局限性的显著挑战。本文提出CVT-Occ,一种新颖方法,通过利用体素在时间上的几何对应关系实现时序融合,以提升3D占位预测精度。通过沿每个体素视线方向采样点,并融合这些点在历史帧中的特征,构建代价体特征图,从而优化当前体素特征,改善预测结果。该方法充分利用历史观测中的视差线索,并采用数据驱动方式学习代价体。我们在Occ3D-Waymo数据集上通过严格实验验证了CVT-Occ的有效性,其在3D占位预测上优于现有最先进方法,且额外计算成本极低。代码已开源。

原文摘要 · Abstract (English)

Vision-based 3D occupancy prediction is significantly challenged by the inherent limitations of monocular vision in depth estimation. This paper introduces CVT-Occ, a novel approach that leverages temporal fusion through the geometric correspondence of voxels over time to improve the accuracy of 3D occupancy predictions. By sampling points along the line of sight of each voxel and integrating the features of these points from historical frames, we construct a cost volume feature map that refines current volume features for improved prediction outcomes. Our method takes advantage of parallax cues from historical observations and employs a data-driven approach to learn the cost volume. We validate the effectiveness of CVT-Occ through rigorous experiments on the Occ3D-Waymo dataset, where it outperforms state-of-the-art methods in 3D occupancy prediction with minimal additional computational cost. The code is released at \url{https://github.com/Tsinghua-MARS-Lab/CVT-Occ}.

3D占位时序融合单目视觉自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。