arXiv:2511.12432cs.CV2025-11AAAI被引 11

通过通道扰动与预训练知识融合,提升多模态图像融合性能。

Text-Guided Channel Perturbation and Pretrained Knowledge Integration for Unified Multi-Modality Image Fusion

  • 引入语义感知通道剪枝与几何仿射调制模块,增强特征区分性。
  • 在多个数据集上超越现有方法,下游任务性能提升显著。
  • 适合多模态图像融合研究者与跨模态视觉任务开发者。

多模态图像融合通过整合互补信息提升场景感知能力。统一模型旨在跨模态共享参数以实现融合,但模态间差异常引发梯度冲突,限制性能。部分方法采用模态专用编码器增强特征感知,改善融合质量,但降低跨任务泛化能力。为此,本文提出基于通道扰动与预训练知识融合的统一多模态图像融合框架(UP-Fusion)。为抑制冗余模态信息并强调关键特征,提出语义感知通道剪枝模块(SCPM),利用预训练模型的语义感知能力筛选和增强多模态特征通道。进一步提出几何仿射调制模块(GAM),利用原始模态特征对初始融合特征施加仿射变换,保持特征编码器的模态判别性。最后,在解码阶段引入文本引导通道扰动模块(TCPM),重塑通道分布,降低对模态特异性通道的依赖。大量实验表明,所提算法在多模态图像融合及下游任务上均优于现有方法。

原文摘要 · Abstract (English)

Multi-modality image fusion enhances scene perception by combining complementary information. Unified models aim to share parameters across modalities for multi-modality image fusion, but large modality differences often cause gradient conflicts, limiting performance. Some methods introduce modality-specific encoders to enhance feature perception and improve fusion quality. However, this strategy reduces generalisation across different fusion tasks. To overcome this limitation, we propose a unified multi-modality image fusion framework based on channel perturbation and pre-trained knowledge integration (UP-Fusion). To suppress redundant modal information and emphasize key features, we propose the Semantic-Aware Channel Pruning Module (SCPM), which leverages the semantic perception capability of a pre-trained model to filter and enhance multi-modality feature channels. Furthermore, we proposed the Geometric Affine Modulation Module (GAM), which uses original modal features to apply affine transformations on initial fusion features to maintain the feature encoder modal discriminability. Finally, we apply a Text-Guided Channel Perturbation Module (TCPM) during decoding to reshape the channel distribution, reducing the dependence on modality-specific channels. Extensive experiments demonstrate that the proposed algorithm outperforms existing methods on both multi-modality image fusion and downstream tasks.

多模态融合通道剪枝预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。