arXiv:2603.21820cs.CV2026-03中稿 · CVPR

打破红外与可见光图像配对训练限制,实现低数据成本下的高性能融合

Beyond Strict Pairing: Arbitrarily Paired Training for High-Performance Infrared and Visible Image Fusion

  • 提出任意配对训练框架,无需严格图像对齐即可学习跨模态关系
  • 在仅1%标注数据下性能接近传统方法100倍数据量的效果
  • 适用于数据稀缺或难以对齐的红外-可见光融合场景

红外与可见光图像融合(IVIF)通过结合互补模态信息,在保留自然纹理和显著热特征的同时实现更优视觉感知。现有方法依赖大量严格对齐的图像对进行训练,但实际获取这类数据成本高、耗时长。此外,训练中强制配对限制了跨模态关系的多样性,影响模型泛化能力。为此,本文挑战严格配对训练范式(SPTP),系统研究无配对(UPTP)与任意配对(APTP)训练策略。建立APTP的理论目标,揭示其与SPTP的互补性;设计可实用的框架,在极有限且未对齐的数据下显著丰富跨模态关系。通过三类轻量级端到端基线(CNN、Transformer、GAN)及创新损失函数验证,即使在内容不一致、数据量仅为常规1%的条件下,仍能实现与传统100×更大数据集上SPTP相当的性能。该成果大幅降低数据采集成本,提升模型鲁棒性,为IVIF研究提供可行新路径。代码已开源。

原文摘要 · Abstract (English)

Infrared and visible image fusion(IVIF) combines complementary modalities while preserving natural textures and salient thermal signatures. Existing solutions predominantly rely on extensive sets of rigidly aligned image pairs for training. However, acquiring such data is often impractical due to the costly and labour-intensive alignment process. Besides, maintaining a rigid pairing setting during training restricts the volume of cross-modal relationships, thereby limiting generalisation performance. To this end, this work challenges the necessity of Strictly Paired Training Paradigm (SPTP) by systematically investigating UnPaired and Arbitrarily Paired Training Paradigms (UPTP and APTP) for high-performance IVIF. We establish a theoretical objective of APTP, reflecting the complementary nature between UPTP and SPTP. More importantly, we develop a practical framework capable of significantly enriching cross-modal relationships even with severely limited and unaligned training data. To validate our propositions, three end-to-end lightweight baselines, alongside a set of innovative loss functions, are designed to cover three classic frameworks (CNN, Transformer, GAN). Comprehensive experiments demonstrate that the proposed APTP and UPTP are feasible and capable of training models on a severely limited and content-inconsistent infrared and visible dataset, achieving performance comparable to that of a dataset 100$\times$ larger in SPTP. This finding fundamentally alleviates the cost and difficulty of data collection while enhancing model robustness from the data perspective, delivering a feasible solution for IVIF studies. The code is available at \href{https://github.com/yanglinDeng/IVIF_unpair}{\textcolor{blue}{https://github.com/yanglinDeng/IVIF\_unpair}}.

图像融合跨模态无监督训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。