arXiv:2509.09085cs.CV2025-09被引 4

通过迭代差分反馈机制,提升多光谱目标检测的特征融合质量。

IRDFusion: Iterative Relation-Map Difference guided Feature Fusion for Multispectral Object Detection

  • 基于跨模态对比与筛选策略,动态融合互补特征。
  • 在FLIR、LLVIP和M^3FD上均达当前最佳性能。
  • 适合多光谱图像分析、红外可见融合等场景使用。

现有多光谱目标检测方法在特征融合时常保留冗余背景或噪声,影响感知性能。为此,本文提出一种创新的特征融合框架,基于跨模态特征对比与筛选策略,区别于传统方法。该方法通过融合对象感知的互补跨模态特征,自适应增强显著结构,同时抑制共享背景干扰。核心包含两个新设计模块:互特征精炼模块(MFRM)与差分特征反馈模块(DFFM)。MFRM通过建模跨模态特征间关系,提升模态内与模态间表示能力,改善对齐性与判别力;受反馈差分放大器启发,DFFM动态计算跨模态差分特征作为引导信号,反馈至MFRM,实现互补信息的自适应融合与跨模态共模噪声抑制。二者集成形成统一框架,正式定义为迭代关系图差分引导特征融合机制(IRDFusion),通过迭代反馈逐步放大显著关系信号,抑制特征噪声,显著提升融合质量。在FLIR、LLVIP和M^3FD数据集上的大量实验表明,IRDFusion在多种复杂场景下持续优于现有方法,验证了其鲁棒性与有效性。代码将开源于https://github.com/61s61min/IRDFusion.git。

原文摘要 · Abstract (English)

Current multispectral object detection methods often retain extraneous background or noise during feature fusion, limiting perceptual performance. To address this, we propose an innovative feature fusion framework based on cross-modal feature contrastive and screening strategy, diverging from conventional approaches. The proposed method adaptively enhances salient structures by fusing object-aware complementary cross-modal features while suppressing shared background interference. Our solution centers on two novel, specially designed modules: the Mutual Feature Refinement Module (MFRM) and the Differential Feature Feedback Module (DFFM). The MFRM enhances intra- and inter-modal feature representations by modeling their relationships, thereby improving cross-modal alignment and discriminative power. Inspired by feedback differential amplifiers, the DFFM dynamically computes inter-modal differential features as guidance signals and feeds them back to the MFRM, enabling adaptive fusion of complementary information while suppressing common-mode noise across modalities. To enable robust feature learning, the MFRM and DFFM are integrated into a unified framework, which is formally formulated as an Iterative Relation-Map Differential Guided Feature Fusion mechanism, termed IRDFusion. IRDFusion enables high-quality cross-modal fusion by progressively amplifying salient relational signals through iterative feedback, while suppressing feature noise, leading to significant performance gains. In extensive experiments on FLIR, LLVIP and M$^3$FD datasets, IRDFusion achieves state-of-the-art performance and consistently outperforms existing methods across diverse challenging scenarios, demonstrating its robustness and effectiveness. Code will be available at https://github.com/61s61min/IRDFusion.git.

多光谱检测特征融合红外可见融合目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。