arXiv:2503.15211cs.CV2025-03CVPR被引 9

用神经辐射场优化多视角3D检测,提升遮挡下物体定位精度

GO-N3RDet: Geometry Optimized NeRF-enhanced 3D Object Detector

  • 通过嵌入3D位置信息优化体素,融合多视角特征
  • 在ScanNet和ARKITScenes上达到新最好性能
  • 适合做多视角3D检测的开发者参考

我们提出GO-N3RDet,一种基于神经辐射场(NeRF)增强的场景几何优化多视角3D目标检测器。准确的3D目标检测依赖于有效的体素表示,但因遮挡和缺乏3D信息,从多视角2D图像构建3D特征极具挑战。为此,我们设计了一种嵌入3D位置信息的体素优化机制以融合多视角特征。为优先重建目标区域的神经场,引入双重要性采样策略。此外,提出不透明度优化模块,通过多视角一致性约束实现精确体素不透明度预测。为进一步提升多视角间体素密度一致性,引入射线距离作为权重因子,减少累积射线误差。这些模块协同构成端到端神经模型,在ScanNet和ARKITScenes数据集上验证了其优越性。代码将开源。

原文摘要 · Abstract (English)

We propose GO-N3RDet, a scene-geometry optimized multi-view 3D object detector enhanced by neural radiance fields. The key to accurate 3D object detection is in effective voxel representation. However, due to occlusion and lack of 3D information, constructing 3D features from multi-view 2D images is challenging. Addressing that, we introduce a unique 3D positional information embedded voxel optimization mechanism to fuse multi-view features. To prioritize neural field reconstruction in object regions, we also devise a double importance sampling scheme for the NeRF branch of our detector. We additionally propose an opacity optimization module for precise voxel opacity prediction by enforcing multi-view consistency constraints. Moreover, to further improve voxel density consistency across multiple perspectives, we incorporate ray distance as a weighting factor to minimize cumulative ray errors. Our unique modules synergetically form an end-to-end neural model that establishes new state-of-the-art in NeRF-based multi-view 3D detection, verified with extensive experiments on ScanNet and ARKITScenes. Code will be available at https://github.com/ZechuanLi/GO-N3RDet.

3D检测神经辐射场多视角体素优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。