arXiv:2411.14716cs.CVcs.LG2024-11CVPR被引 21

用图像自监督训练自动驾驶视觉模型,提升3D感知能力。

VisionPAD: A Vision-Centric Pre-training Paradigm for Autonomous Driving

  • 用3D高斯点云重建多视角图像,仅依赖图像监督
  • 通过体素运动估计和多帧光度一致性,学习序列运动与几何信息
  • 在检测、占位预测等任务上超越现有预训练方法

本文提出VisionPAD,一种面向自动驾驶视觉算法的新型自监督预训练范式。与以往依赖显式深度监督的神经渲染方法不同,VisionPAD采用更高效的3D高斯点云技术,仅使用图像作为监督信号重建多视角表征。具体地,提出一种自监督体素速度估计方法:通过将体素映射到相邻帧并监督渲染输出,模型有效学习序列数据中的运动线索。此外,引入多帧光度一致性策略,基于渲染深度和相对位姿将相邻帧投影至当前帧,通过纯图像监督增强三维几何感知。在自动驾驶数据集上的大量实验表明,VisionPAD显著提升3D目标检测、占位预测和地图分割性能,明显优于现有最先进的预训练策略。

原文摘要 · Abstract (English)

This paper introduces VisionPAD, a novel self-supervised pre-training paradigm designed for vision-centric algorithms in autonomous driving. In contrast to previous approaches that employ neural rendering with explicit depth supervision, VisionPAD utilizes more efficient 3D Gaussian Splatting to reconstruct multi-view representations using only images as supervision. Specifically, we introduce a self-supervised method for voxel velocity estimation. By warping voxels to adjacent frames and supervising the rendered outputs, the model effectively learns motion cues in the sequential data. Furthermore, we adopt a multi-frame photometric consistency approach to enhance geometric perception. It projects adjacent frames to the current frame based on rendered depths and relative poses, boosting the 3D geometric representation through pure image supervision. Extensive experiments on autonomous driving datasets demonstrate that VisionPAD significantly improves performance in 3D object detection, occupancy prediction and map segmentation, surpassing state-of-the-art pre-training strategies by a considerable margin.

自动驾驶自监督3D感知视觉预训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。