arXiv:2506.19615cs.CV2025-06

自监督多模态NeRF实现自动驾驶场景的时空联合建模

Self-Supervised Multimodal NeRF for Autonomous Driving

  • 基于自监督学习,联合建模激光雷达与摄像头的时空动态场景
  • 在KITTI-360上优于基线模型,激光雷达与相机域均表现最佳
  • 适合自动驾驶、三维场景重建方向的研究者参考

本文提出一种基于神经辐射场(NeRF)的框架——新视角合成框架(NVSF),联合学习静态与动态场景中激光雷达和摄像头的隐式时空表示。该方法在真实自动驾驶场景下进行测试,具有自监督特性,无需3D标注数据。为提升训练效率并加速收敛,引入基于启发式的图像像素采样策略,聚焦高信息量像素;为保留激光雷达点的局部特征,采用双梯度掩码机制。在KITTI-360数据集上的大量实验表明,相较于基线模型,本框架在激光雷达与相机两个模态上均取得最优性能。模型代码已开源。

原文摘要 · Abstract (English)

In this paper, we propose a Neural Radiance Fields (NeRF) based framework, referred to as Novel View Synthesis Framework (NVSF). It jointly learns the implicit neural representation of space and time-varying scene for both LiDAR and Camera. We test this on a real-world autonomous driving scenario containing both static and dynamic scenes. Compared to existing multimodal dynamic NeRFs, our framework is self-supervised, thus eliminating the need for 3D labels. For efficient training and faster convergence, we introduce heuristic-based image pixel sampling to focus on pixels with rich information. To preserve the local features of LiDAR points, a Double Gradient based mask is employed. Extensive experiments on the KITTI-360 dataset show that, compared to the baseline models, our framework has reported best performance on both LiDAR and Camera domain. Code of the model is available at https://github.com/gaurav00700/Selfsupervised-NVSF

NeRF自动驾驶多模态自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。