arXiv:2410.15814cs.CVcs.AI2024-10被引 3

用KAN网络提升路侧摄像头与激光雷达3D感知融合效果

Kaninfradet3D:A Road-side Camera-LiDAR Fusion 3D Perception Model based on Nonlinear Feature Extraction and Intrinsic Correlation

  • 引入KAN网络替代MLP,更好提取摄像头和激光雷达的高维特征
  • 在TUMTraf数据集上提升9.87~10.64 mAP,路侧视角提升1.40 mAP
  • 适合关注路侧协同感知、多模态融合的自动驾驶研究者

随着智能驾驶发展,车辆自身3D感知方法众多,但路侧感知研究较少。路侧视角具备全局视野和更广感知范围,具有重要价值。激光雷达提供精确三维空间信息,摄像头则富含语义信息,二者互补性强。然而,现有方法中加入摄像头数据并未提升精度,因特征提取与融合机制不可靠。近期提出的柯尔莫哥洛夫-阿诺德网络(KAN)可替代传统MLP,更适合处理高维复杂数据。本文提出Kaninfradet3D模型,优化特征提取与融合模块:采用KAN层改进编码器与融合器,引入跨注意力机制增强特征融合,可视化显示摄像头特征分布更均匀,解决了特征异常集中问题。在TUMTraf Intersection Dataset两个视角下分别实现+9.87 mAP和+10.64 mAP提升,在TUMTraf V2X Cooperative Perception Dataset路侧端提升+1.40 mAP。结果表明,Kaninfradet3D能有效融合多模态特征,验证了KAN在路侧感知任务中的潜力。

原文摘要 · Abstract (English)

With the development of AI-assisted driving, numerous methods have emerged for ego-vehicle 3D perception tasks, but there has been limited research on roadside perception. With its ability to provide a global view and a broader sensing range, the roadside perspective is worth developing. LiDAR provides precise three-dimensional spatial information, while cameras offer semantic information. These two modalities are complementary in 3D detection. However, adding camera data does not increase accuracy in some studies since the information extraction and fusion procedure is not sufficiently reliable. Recently, Kolmogorov-Arnold Networks (KANs) have been proposed as replacements for MLPs, which are better suited for high-dimensional, complex data. Both the camera and the LiDAR provide high-dimensional information, and employing KANs should enhance the extraction of valuable features to produce better fusion outcomes. This paper proposes Kaninfradet3D, which optimizes the feature extraction and fusion modules. To extract features from complex high-dimensional data, the model's encoder and fuser modules were improved using KAN Layers. Cross-attention was applied to enhance feature fusion, and visual comparisons verified that camera features were more evenly integrated. This addressed the issue of camera features being abnormally concentrated, negatively impacting fusion. Compared to the benchmark, our approach shows improvements of +9.87 mAP and +10.64 mAP in the two viewpoints of the TUMTraf Intersection Dataset and an improvement of +1.40 mAP in the roadside end of the TUMTraf V2X Cooperative Perception Dataset. The results indicate that Kaninfradet3D can effectively fuse features, demonstrating the potential of applying KANs in roadside perception tasks.

3D感知多模态融合KAN路侧感知

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。