基于熵引导的自适应融合,提升无人机红外与可见光目标检测性能
EGM-Det: Entropy-Guided Multimodal Adaptive Fusion for UAV RGB-IR Object Detection

- 用熵偏移门机制动态融合多模态特征,避免固定权重融合
- 在VEDAI数据集上超越前人方法超10个百分点,三数据集均达顶尖水平
- 适合需要高鲁棒性多模态感知的无人机检测场景
联合使用可见光(RGB)与红外(IR)图像可提升无人机视角下的目标检测效果,但现有方法多采用静态或固定权重融合多模态特征,忽视了不同空间位置下模态可靠性差异。本文提出EGM-Det,一种熵引导的多模态自适应融合框架,用于RGB-IR目标检测。该框架采用双流结构保留各模态特异性表示,并引入熵偏移门融合模块,从输入强度、局部熵和跨模态差异中提取浅层熵先验,指导局部偏移对齐与空间-通道门控融合,从而选择性聚合可靠的RGB与红外线索,而非均匀混合异质特征。此外,通过跨模态蒸馏正则化学习到的融合门控,减少融合退化问题。每个学生分支从匹配主干的跨模态教师分支中提取互补知识,熵自适应监督强化对不确定模态决策的关注。在DroneVehicle、LLVIP和VEDAI三个数据集上的实验表明,EGM-Det在所有基准上均达到当前最优性能,尤其在VEDAI上较之前方法提升超过10个百分点。
原文摘要 · Abstract (English)
Joint use of RGB and infrared (IR) imagery can improve UAV-view object detection, but most existing methods fuse multimodal features with static or fixed weights and therefore overlook spatially varying modality reliability. We propose EGM-Det, an entropy-guided multimodal adaptive fusion framework for RGB-IR object detection. EGM-Det employs a dual-stream architecture to preserve modality-specific representations and introduces an Entropy Offset Gate Fusion module for adaptive multi-scale fusion. The module derives shallow entropy priors from input intensity, local entropy, and cross-modal discrepancy, and uses them to guide local offset alignment and spatial-channel gated fusion. It therefore selectively aggregates reliable RGB and infrared cues instead of uniformly combining heterogeneous features. We further introduce cross-modal distillation to regularize the learned fusion gates and reduce fusion degradation. Each student branch extracts complementary knowledge from the cross-modality teacher branch matched to the main branch, while entropy-adaptive supervision emphasizes uncertain modality decisions. Experiments on DroneVehicle, LLVIP, and VEDAI demonstrate state-of-the-art performance across all three benchmarks; in particular, EGM-Det outperforms prior approaches by more than 10 percentage points on VEDAI.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。