arXiv:2609.03516cs.CV2026-09

解决红外可见光目标检测中模态缺失时的融合难题

Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection

论文配图:Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection
图 1 · 摘自论文原文
  • 设计自适应专家协作机制,动态选择跨模态或单模态融合路径
  • 在模态缺失情况下仍保持检测性能,完整与缺失场景均稳定
  • 基于残差自调速最优传输对齐异质特征,提升语义匹配可靠性

红外-可见光目标检测(IVOD)通过融合可见光与红外传感器的互补信息,在复杂场景中实现可靠感知。然而实际中传感器可能失效或丢帧,导致某一模态缺失或间歇性中断。现有方法假设双模态始终存在,固定融合策略在模态缺失时会崩溃。此外,如何在光谱分布差异下准确估计异质模态间的语义关联仍是关键挑战。本文提出FlexibleFusion,一种统一且自适应的方法,可灵活分配融合路径与强度,在完整与缺失模态场景下无缝运行。其核心为模态感知专家协作(MAEC)机制,可选择性激活跨模态或单模态专家路径:双模态时进行跨模态融合,单模态时转为自融合。同时,设计残差自调速熵最优传输(RSPEOT),从传输视角对齐异质特征分布。相比标准熵最优传输(EOT)依赖固定稀疏系数,RSPEOT引入残差驱动的自调速更新,优先保留可靠匹配并逐步优化困难样本。该设计减轻了标准EOT的额外优化负担,同时保持可靠的语义对齐。在完整与缺失模态协议下的综合实验表明,该方法在任意模态配置下均表现一致稳定。代码将在发表后公开。

原文摘要 · Abstract (English)

Infrared-visible object detection (IVOD) integrates complementary evidence from visible and infrared sensors for reliable perception in challenging scenes. In practice, sensors may fail or drop frames, leaving one modality unavailable or intermittent. Existing methods for IVOD assume both modalities are always present, and fixed fusion collapses when one stream is missing. Furthermore, it remains a critical challenge to reliably estimate semantic correlation across heterogeneous modalities, especially under spectral distribution discrepancy. We present FlexibleFusion, a unified and adaptive method that flexibly allocates integration pathways and fusion strength, operating seamlessly across complete and missing-modality regimes. At its core, the Modality-Aware Experts Collaboration (MAEC) mechanism selectively activates and aggregates cross-modal or intra-modal expert pathways. It allows cross-modal fusion when full modalities are available and falls back to self-fusion under missing conditions. Additionally, we design Residual Self-Paced Entropic Optimal Transport (RSPEOT) to align heterogeneous feature distributions from a transport perspective. Instead of relying on the fixed sparsity coefficient in standard entropic optimal transport (EOT), RSPEOT introduces a residual-driven self-paced update that prioritizes reliable matches and progressively refines harder ones. This design alleviates the additional optimization burden of standard EOT while preserving reliable semantic alignment. Comprehensive experiments under complete and missing-modality protocols show consistent performance across arbitrary modality configurations. Code will be released upon publication.

目标检测多模态融合最优传输红外可见光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。