arXiv:2603.28090cs.CV2026-03被引 1

提出新模型避免视图变换冲突,提升3D感知连续表示能力。

To View Transform or Not to View Transform: NeRF-based Pre-training Perspective

  • 用点云构建连续3D表示,避开视图变换的离散限制
  • 在nuScenes上检测精度超越现有最佳方法,重建与检测双优
  • 保留预训练NeRF网络,适配下游任务,提升模型复用效率

神经辐射场(NeRF)已成为视觉主导自动驾驶中一种重要的自监督预训练范式,能有效增强对三维几何与外观的理解。现有方法将NeRF直接应用于经视图变换获得的体素特征,但视图变换采用离散刚性表示,而NeRF假设连续自适应函数,两者矛盾导致3D表征模糊不清,限制了场景理解能力。此外,预训练中的NeRF网络在下游任务中被丢弃,未能充分利用其生成的3D表示。本文提出一种新型点基NeRF类3D检测器(NeRP3D),可学习连续3D表示,避免视图变换带来的先验冲突。该模型始终保留预训练的NeRF网络,继承连续表示学习原则,在场景重建与目标检测任务中均展现更强潜力。在nuScenes数据集上的实验表明,所提方法显著优于现有最先进方法,不仅在预训练重建任务中表现更佳,且在下游检测任务中也实现性能突破。

原文摘要 · Abstract (English)

Neural radiance fields (NeRFs) have emerged as a prominent pre-training paradigm for vision-centric autonomous driving, which enhances 3D geometry and appearance understanding in a fully self-supervised manner. To apply NeRF-based pretraining to 3D perception models, recent approaches have simply applied NeRFs to volumetric features obtained from view transformation. However, coupling NeRFs with view transformation inherits conflicting priors; view transformation imposes discrete and rigid representations, whereas radiance fields assume continuous and adaptive functions. When these opposing assumptions are forced into a single pipeline, the misalignment surfaces as blurry and ambiguous 3D representations that ultimately limit 3D scene understanding. Moreover, the NeRF network for pre-training is discarded during downstream tasks, resulting in inefficient utilization of enhanced 3D representations through NeRF. In this paper, we propose a novel NeRF-Resembled Point-based 3D detector that can learn continuous 3D representation and thus avoid the misaligned priors from view transformation. NeRP3D preserves the pre-trained NeRF network regardless of the tasks, inheriting the principle of continuous 3D representation learning and leading to greater potentials for both scene reconstruction and detection tasks. Experiments on nuScenes dataset demonstrate that our proposed approach significantly improves previous state-of-the-art methods, outperforming not only pretext scene reconstruction tasks but also downstream detection tasks.

3D检测NeRF自动驾驶点云建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。