arXiv:2512.20921cs.CV2025-12AAAI被引 5

提出自监督多模态共识Mamba框架,提升图像融合质量与下游任务表现。

Self-supervised Multiplex Consensus Mamba for General Image Fusion

  • 通过自适应门控与跨模态扫描增强特征表达
  • 在红外-可见光、医学等多场景超越主流方法
  • 无需额外计算开销,适合多任务图像融合应用

图像融合通过整合多模态互补信息生成高质量融合图像,从而提升目标检测、语义分割等下游任务性能。现有方法多针对特定任务,而通用图像融合需在不增加复杂度的前提下兼顾多种任务。为此,我们提出自监督多模态共识Mamba(SMC-Mamba)框架。其中,模态无关特征增强(MAFE)模块通过自适应门控保留细节,并利用空间-通道与频率旋转扫描增强全局表征;多路共识跨模态Mamba(MCCM)模块实现专家间动态协作并达成共识,高效融合多模态互补信息,其跨模态扫描进一步强化模态间特征交互。此外,引入双层自监督对比学习损失(BSCL),在不增加计算开销的前提下保留高频信息,同时提升下游任务性能。大量实验表明,该方法在红外-可见光、医学、多焦点、多曝光融合以及下游视觉任务中均优于当前最先进(SOTA)算法。

原文摘要 · Abstract (English)

Image fusion integrates complementary information from different modalities to generate high-quality fused images, thereby enhancing downstream tasks such as object detection and semantic segmentation. Unlike task-specific techniques that primarily focus on consolidating inter-modal information, general image fusion needs to address a wide range of tasks while improving performance without increasing complexity. To achieve this, we propose SMC-Mamba, a Self-supervised Multiplex Consensus Mamba framework for general image fusion. Specifically, the Modality-Agnostic Feature Enhancement (MAFE) module preserves fine details through adaptive gating and enhances global representations via spatial-channel and frequency-rotational scanning. The Multiplex Consensus Cross-modal Mamba (MCCM) module enables dynamic collaboration among experts, reaching a consensus to efficiently integrate complementary information from multiple modalities. The cross-modal scanning within MCCM further strengthens feature interactions across modalities, facilitating seamless integration of critical information from both sources. Additionally, we introduce a Bi-level Self-supervised Contrastive Learning Loss (BSCL), which preserves high-frequency information without increasing computational overhead while simultaneously boosting performance in downstream tasks. Extensive experiments demonstrate that our approach outperforms state-of-the-art (SOTA) image fusion algorithms in tasks such as infrared-visible, medical, multi-focus, and multi-exposure fusion, as well as downstream visual tasks.

图像融合自监督学习多模态Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。