提出一种通用多模态融合架构,提升红外与可见光目标检测性能。
Towards a Generalizable Fusion Architecture for Multimodal Object Detection
- 通过频域滤波和交叉注意力机制融合红外与可见光特征。
- 在VEDAI和LLVIP数据集上分别提升13.9%和1.1%的mAP@50。
- 无需特定数据集调优,适用于多种复杂场景下的多模态检测。
多模态目标检测通过融合来自不同传感器的互补信息,在挑战性环境下提升鲁棒性。本文提出过滤式多模态交叉注意力融合(FMCAF),一种用于增强RGB与红外(IR)输入融合的预处理架构。FMCAF结合频域滤波模块(Freq-Filter)以抑制冗余光谱特征,并引入基于交叉注意力的融合模块(MCAF)以促进模态间特征共享。与针对特定数据集设计的方法不同,FMCAF旨在实现通用性,在无需数据集特异性调优的情况下,提升多种多模态检测任务的表现。在LLVIP(低光行人检测)和VEDAI(航拍车辆检测)数据集上,FMCAF优于传统融合方法(拼接),分别实现+13.9% mAP@50和+1.1% mAP@50的提升。结果表明,FMCAF可作为未来鲁棒多模态检测流水线的灵活基础。
原文摘要 · Abstract (English)
Multimodal object detection improves robustness in chal- lenging conditions by leveraging complementary cues from multiple sensor modalities. We introduce Filtered Multi- Modal Cross Attention Fusion (FMCAF), a preprocess- ing architecture designed to enhance the fusion of RGB and infrared (IR) inputs. FMCAF combines a frequency- domain filtering block (Freq-Filter) to suppress redun- dant spectral features with a cross-attention-based fusion module (MCAF) to improve intermodal feature sharing. Unlike approaches tailored to specific datasets, FMCAF aims for generalizability, improving performance across different multimodal challenges without requiring dataset- specific tuning. On LLVIP (low-light pedestrian detec- tion) and VEDAI (aerial vehicle detection), FMCAF outper- forms traditional fusion (concatenation), achieving +13.9% mAP@50 on VEDAI and +1.1% on LLVIP. These results support the potential of FMCAF as a flexible foundation for robust multimodal fusion in future detection pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。