arXiv:2507.13899cs.CV2025-07被引 3

用视觉模型深度先验增强激光点特征,提升3D目标检测精度

Enhancing LiDAR Point Features with Foundation Model Priors for 3D Object Detection

  • 引入DepthAnything的深度先验,融合到原始激光点特征中
  • 在KITTI上实现更优检测性能,显著提升小物体识别能力
  • 适合自动驾驶中依赖激光雷达的3D感知任务

近期基础模型的发展为提升三维感知提供了新可能。尤其是从单目图像中获得密集且可靠的几何先验的DepthAnything,可弥补自动驾驶场景下稀疏激光雷达数据的不足。然而,这些先验在基于激光雷达的3D目标检测中仍未被充分使用。本文针对原始激光点特征表达能力有限、特别是反射率属性区分度弱的问题,引入由DepthAnything预测的深度先验,并将其与原始激光属性融合,以丰富每个点的表征。为有效利用增强后的点特征,我们提出一种点级特征提取模块;进一步采用双路径感兴趣区域(RoI)特征提取框架,包含基于体素的分支用于全局语义上下文,以及基于点的分支用于细粒度结构细节。为有效融合互补的RoI特征,设计了双向门控的RoI特征融合模块,平衡全局与局部线索。在KITTI基准上的大量实验表明,该方法持续提升检测准确率,验证了将视觉基础模型先验引入激光雷达3D目标检测的有效性。

原文摘要 · Abstract (English)

Recent advances in foundation models have opened up new possibilities for enhancing 3D perception. In particular, DepthAnything offers dense and reliable geometric priors from monocular RGB images, which can complement sparse LiDAR data in autonomous driving scenarios. However, such priors remain underutilized in LiDAR-based 3D object detection. In this paper, we address the limited expressiveness of raw LiDAR point features, especially the weak discriminative capability of the reflectance attribute, by introducing depth priors predicted by DepthAnything. These priors are fused with the original LiDAR attributes to enrich each point's representation. To leverage the enhanced point features, we propose a point-wise feature extraction module. Then, a Dual-Path RoI feature extraction framework is employed, comprising a voxel-based branch for global semantic context and a point-based branch for fine-grained structural details. To effectively integrate the complementary RoI features, we introduce a bidirectional gated RoI feature fusion module that balances global and local cues. Extensive experiments on the KITTI benchmark show that our method consistently improves detection accuracy, demonstrating the value of incorporating visual foundation model priors into LiDAR-based 3D object detection.

3D检测激光雷达深度先验多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。