arXiv:2608.02092cs.CVcs.MM2026-08

通过空间掩码与通道竞争机制,提升多模态目标检测的鲁棒性。

Deep Multimodal Fusion Detection through Spatial Mask and Channel Competition

论文配图:Deep Multimodal Fusion Detection through Spatial Mask and Channel Competition
图 1 · 摘自论文原文
  • 引入语义掩码交换,训练时主动混合模态边界以减少依赖。
  • 设计可学习通道竞争机制,实现通道级特征采样与聚合。
  • 在多个数据集上表现优异,适合多模态检测任务研究者参考。

多模态目标检测通过挖掘模态特性已展现良好性能。然而,现有特征级融合方法主要在双模态间加权并统一于同一表征空间,易导致双主干架构中单一模态统计特性的过拟合或过度专化。本文提出一种注意力驱动的互补重采样框架,以增强跨模态目标检测的鲁棒性。基于共享通道空间注意力机制,首先引入语义掩码交换,在训练阶段主动混合模态边界,迫使主干网络学习不依赖固定模态标签的通用特征。随后提出可学习通道竞争机制,以通道级且可学习的方式采样并聚合特征。在多个数据集上的实验表明,所提方法有效,结果达到现有先进水平。源代码见补充材料。

原文摘要 · Abstract (English)

Deep multimodal fusion for object detection has demonstrated good performance through mining modal characteristics. However, existing feature-level fusion methods mainly weigh between two modalities and unify them in a unified representation space. This can lead to overfitting or over-specialization of the statistical properties of a single modality within a dual-backbone architecture. This paper proposes an Attention-Driven Complementarity Resampling framework for robust improvement of cross-modality object detection. Based on a shared channel spatial attention mechanism, we first introduce the semantic mask exchange to actively mix the boundaries of the modalities during the training phase, forcing the backbone network to learn generalized features without relying on fixed modal labels. Then we propose a learnable channel competition to sample and aggregate features in a channel-wise and learnable way. Our experiments on multiple datasets demonstrate that the proposed method is effective and yields competitive results among existing state-of-the-art approaches. The source code is provided in the supplementary material.

多模态融合目标检测注意力机制特征融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。