arXiv:2504.03563cs.CV2025-04CVPR被引 8

用提示词增强多模态特征融合,小样本下3D检测性能提升

PF3Det: A Prompted Foundation Feature Assisted Visual LiDAR 3D Detector

  • 引入基础模型与软提示,优化激光雷达与图像特征融合
  • 在nuScenes上仅用少量数据,NDS提升1.19%,mAP提升2.42%
  • 适合资源受限场景下的自动驾驶感知系统研发

3D目标检测对自动驾驶至关重要,需结合激光雷达点云的精确深度信息和相机图像的丰富语义信息。多模态方法虽能提升检测鲁棒性,但跨模态特征融合仍受域差异影响。此外,模型性能常受限于高质量标注数据稀缺,而人工标注成本高昂。近期基础模型在多模态大规模预训练方面取得进展,结合提示工程可实现高效训练。本文提出提示引导的基础3D检测器(PF3Det),融合基础模型编码器与软提示,强化激光雷达-相机特征融合能力。在nuScenes数据集上,该方法在有限训练数据下达到当前最优性能,NDS提升1.19%,mAP提升2.42%,验证了其在3D检测中的高效性。

原文摘要 · Abstract (English)

3D object detection is crucial for autonomous driving, leveraging both LiDAR point clouds for precise depth information and camera images for rich semantic information. Therefore, the multi-modal methods that combine both modalities offer more robust detection results. However, efficiently fusing LiDAR points and images remains challenging due to the domain gaps. In addition, the performance of many models is limited by the amount of high quality labeled data, which is expensive to create. The recent advances in foundation models, which use large-scale pre-training on different modalities, enable better multi-modal fusion. Combining the prompt engineering techniques for efficient training, we propose the Prompted Foundational 3D Detector (PF3Det), which integrates foundation model encoders and soft prompts to enhance LiDAR-camera feature fusion. PF3Det achieves the state-of-the-art results under limited training data, improving NDS by 1.19% and mAP by 2.42% on the nuScenes dataset, demonstrating its efficiency in 3D detection.

3D检测多模态融合提示工程自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。