arXiv:2510.15991cs.CV2025-10被引 2

通过几何与分布引导提升稀疏多模态3D检测性能,兼顾精度与速度。

CrossRay3D: Geometry and Distribution Guidance for Efficient Multimodal 3D Detection

  • 设计射线感知监督与类别平衡机制,优化稀疏特征表示质量。
  • 在nuScenes上达到72.4 mAP和74.7 NDS,比领先方法快1.84倍。
  • 对缺失激光雷达或摄像头数据具有强鲁棒性,适合实际部署场景。

稀疏跨模态检测器相比鸟瞰图(BEV)检测器,在下游任务适应性和计算成本方面更具优势。然而现有稀疏检测器忽视了标记项的表示质量,导致前景表征不佳、性能受限。本文指出几何结构保持与类别分布是关键,提出稀疏选择器(SS)。其核心为射线感知监督(RAS),训练中保留丰富几何信息;以及类别平衡监督,自适应重加权类别语义重要性,确保小目标相关标记在采样中被保留。同时引入射线位置编码(Ray PE)以缓解激光雷达与图像之间的分布差异。将上述模块集成至端到端稀疏多模态检测器CrossRay3D。实验表明,在挑战性的nuScenes基准上,CrossRay3D达到72.4 mAP和74.7 NDS,且运行速度比其他先进方法快1.84倍。此外,即使在激光雷达或相机数据部分或完全缺失时,仍表现出强鲁棒性。

原文摘要 · Abstract (English)

The sparse cross-modality detector offers more advantages than its counterpart, the Bird's-Eye-View (BEV) detector, particularly in terms of adaptability for downstream tasks and computational cost savings. However, existing sparse detectors overlook the quality of token representation, leaving it with a sub-optimal foreground quality and limited performance. In this paper, we identify that the geometric structure preserved and the class distribution are the key to improving the performance of the sparse detector, and propose a Sparse Selector (SS). The core module of SS is Ray-Aware Supervision (RAS), which preserves rich geometric information during the training stage, and Class-Balanced Supervision, which adaptively reweights the salience of class semantics, ensuring that tokens associated with small objects are retained during token sampling. Thereby, outperforming other sparse multi-modal detectors in the representation of tokens. Additionally, we design Ray Positional Encoding (Ray PE) to address the distribution differences between the LiDAR modality and the image. Finally, we integrate the aforementioned module into an end-to-end sparse multi-modality detector, dubbed CrossRay3D. Experiments show that, on the challenging nuScenes benchmark, CrossRay3D achieves state-of-the-art performance with 72.4 mAP and 74.7 NDS, while running 1.84 faster than other leading methods. Moreover, CrossRay3D demonstrates strong robustness even in scenarios where LiDAR or camera data are partially or entirely missing.

3D检测多模态稀疏建模自动驾驶

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。