解决多模态检测中单模态特征退化问题,提升检测性能
Rethinking Multi-Modal Object Detection from the Perspective of Mono-Modality Feature Learning
- 提出单模态蒸馏与光照感知融合机制,增强单模态学习能力
- 在三个数据集上超越当前最优模型,缓解特征融合退化现象
- 适合关注多模态检测优化与轻量级融合策略的研究者
多模态目标检测(MMOD)因在复杂环境下的强适应性而广泛应用。现有研究多聚焦于RGB-IR模态间的互补特征融合,却忽视了多模态联合训练导致的单模态特征提取能力下降问题,引发普遍存在的‘融合退化’现象,阻碍性能提升。为此,本文引入线性探针评估,从单模态学习角度重新思考MMOD任务,提出M²D-LIF框架,包含单模态蒸馏(M²D)和局部光照感知融合(LIF)模块。该框架在多模态联合训练中促进单模态充分学习,并实现轻量高效特征融合,显著提升检测性能。在三个MMOD数据集上的实验表明,M²D-LIF有效缓解融合退化,优于现有最先进方法。代码已开源。
原文摘要 · Abstract (English)
Multi-Modal Object Detection (MMOD), due to its stronger adaptability to various complex environments, has been widely applied in various applications. Extensive research is dedicated to the RGB-IR object detection, primarily focusing on how to integrate complementary features from RGB-IR modalities. However, they neglect the mono-modality insufficient learning problem, which arises from decreased feature extraction capability in multi-modal joint learning. This leads to a prevalent but unreasonable phenomenon\textemdash Fusion Degradation, which hinders the performance improvement of the MMOD model. Motivated by this, in this paper, we introduce linear probing evaluation to the multi-modal detectors and rethink the multi-modal object detection task from the mono-modality learning perspective. Therefore, we construct a novel framework called M$^2$D-LIF, which consists of the Mono-Modality Distillation (M$^2$D) method and the Local Illumination-aware Fusion (LIF) module. The M$^2$D-LIF framework facilitates the sufficient learning of mono-modality during multi-modal joint training and explores a lightweight yet effective feature fusion manner to achieve superior object detection performance. Extensive experiments conducted on three MMOD datasets demonstrate that our M$^2$D-LIF effectively mitigates the Fusion Degradation phenomenon and outperforms the previous SOTA detectors. The codes are available at https://github.com/Zhao-Tian-yi/M2D-LIF.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。