自适应融合多模态数据,提升复杂场景下3D目标检测鲁棒性
AG-Fusion: adaptive gated multimodal fusion for 3d object detection in complex scenes
- 通过跨模态注意力动态选择可靠特征进行融合
- 在标准KITTI上达93.92%准确率,在挑战性E3D上提升24.88%
- 专为挖掘机作业场景设计新数据集E3D,适合工业自动驾驶研究
多模态相机-LiDAR融合技术在3D目标检测中应用广泛,表现良好。但现有方法在传感器退化或环境干扰的复杂场景中性能显著下降。本文提出一种新型自适应门控融合(AG-Fusion)方法,通过识别可靠模式实现跨模态知识的选择性整合,提升复杂场景下的检测鲁棒性。首先将各模态特征投影至统一的鸟瞰图(BEV)空间,并使用基于窗口的注意力机制增强;随后设计基于跨模态注意力的自适应门控融合模块,生成对复杂环境鲁棒的BEV表示。此外,构建新数据集Excavator3D(E3D),聚焦于具有挑战性的挖掘机作业场景,用于评估复杂条件下的性能。所提方法在标准KITTI数据集上达到93.92%的准确率,且在挑战性E3D数据集上相比基线提升24.88%,验证了其对不可靠模态信息的强鲁棒性。
原文摘要 · Abstract (English)
Multimodal camera-LiDAR fusion technology has found extensive application in 3D object detection, demonstrating encouraging performance. However, existing methods exhibit significant performance degradation in challenging scenarios characterized by sensor degradation or environmental disturbances. We propose a novel Adaptive Gated Fusion (AG-Fusion) approach that selectively integrates cross-modal knowledge by identifying reliable patterns for robust detection in complex scenes. Specifically, we first project features from each modality into a unified BEV space and enhance them using a window-based attention mechanism. Subsequently, an adaptive gated fusion module based on cross-modal attention is designed to integrate these features into reliable BEV representations robust to challenging environments. Furthermore, we construct a new dataset named Excavator3D (E3D) focusing on challenging excavator operation scenarios to benchmark performance in complex conditions. Our method not only achieves competitive performance on the standard KITTI dataset with 93.92% accuracy, but also significantly outperforms the baseline by 24.88% on the challenging E3D dataset, demonstrating superior robustness to unreliable modal information in complex industrial scenes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。