arXiv:2410.11358cs.CV2024-10被引 16

通过对比学习提升Transformer在多模态检测中的深层语义提取能力

SeaDATE: Remedy Dual-Attention Transformer with Semantic Alignment via Contrast Learning for Multimodal Object Detection

  • 设计双注意力融合模块,从空间与通道双视角增强跨模态特征对齐
  • 在FLIR、LLVIP、M3FD数据集上达到最新性能,显著优于现有方法
  • 适合关注多模态目标检测中深层语义建模的研究者与工程师

多模态目标检测利用多种模态信息提升检测精度与鲁棒性。Transformer通过学习长程依赖,在特征提取阶段有效融合多模态特征,显著提升检测性能。然而,现有方法仅堆叠Transformer引导的融合技术,未充分挖掘网络不同深度层的特征提取能力,限制了性能提升。本文提出一种高效准确的检测方法SeaDATE。首先,设计新型双注意力特征融合(DTF)模块,在Transformer引导下,通过空间与通道令牌从局部和全局角度融合多模态特征。理论分析与实验证明,基于Transformer的融合方法在浅层特征细节上表现更优,而深层语义信息提取不足。为此,我们引入对比学习(CL)模块,学习多模态样本特征,弥补Transformer融合在深层语义特征提取上的短板,有效利用跨模态信息。在FLIR、LLVIP和M3FD数据集上的大量实验与消融研究证明该方法有效,达到当前最优性能。

原文摘要 · Abstract (English)

Multimodal object detection leverages diverse modal information to enhance the accuracy and robustness of detectors. By learning long-term dependencies, Transformer can effectively integrate multimodal features in the feature extraction stage, which greatly improves the performance of multimodal object detection. However, current methods merely stack Transformer-guided fusion techniques without exploring their capability to extract features at various depth layers of network, thus limiting the improvements in detection performance. In this paper, we introduce an accurate and efficient object detection method named SeaDATE. Initially, we propose a novel dual attention Feature Fusion (DTF) module that, under Transformer's guidance, integrates local and global information through a dual attention mechanism, strengthening the fusion of modal features from orthogonal perspectives using spatial and channel tokens. Meanwhile, our theoretical analysis and empirical validation demonstrate that the Transformer-guided fusion method, treating images as sequences of pixels for fusion, performs better on shallow features' detail information compared to deep semantic information. To address this, we designed a contrastive learning (CL) module aimed at learning features of multimodal samples, remedying the shortcomings of Transformer-guided fusion in extracting deep semantic features, and effectively utilizing cross-modal information. Extensive experiments and ablation studies on the FLIR, LLVIP, and M3FD datasets have proven our method to be effective, achieving state-of-the-art detection performance.

多模态检测Transformer对比学习特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。