arXiv:2502.14493cs.CVcs.LG2025-02被引 14

提升红外可见光图像融合在真实场景下的鲁棒性

CrossFuse: Learning Infrared and Visible Image Fusion by Cross-Sensor Top-K Vision Alignment and Beyond

  • 通过跨传感器Top-k视觉对齐实现外部数据增强
  • 自监督学习结合弱-强增强,提升特征泛化能力
  • 显著改善复杂环境下图像融合的稳定性与可靠性

红外与可见光图像融合(IVIF)在视频监控和自动驾驶等关键领域应用日益广泛。尽管基于深度学习的方法已取得显著进展,但实际应用中模型常遭遇分布外(OOD)场景,严重降低性能与可靠性。针对此问题,本文提出基于多视角增强的红外-可见光图像融合框架。外部数据增强采用逐通道的可见光图像变换,通过Top-k选择性视觉对齐缓解数据集间分布偏移;内部数据增强则引入弱-强增强策略构建自监督学习机制,使模型在融合过程中学习更鲁棒、泛化的特征表示。大量实验表明,所提方法在多种条件与环境中均表现优异,显著提升了IVIF任务在实际应用中的可靠性和稳定性。

原文摘要 · Abstract (English)

Infrared and visible image fusion (IVIF) is increasingly applied in critical fields such as video surveillance and autonomous driving systems. Significant progress has been made in deep learning-based fusion methods. However, these models frequently encounter out-of-distribution (OOD) scenes in real-world applications, which severely impact their performance and reliability. Therefore, addressing the challenge of OOD data is crucial for the safe deployment of these models in open-world environments. Unlike existing research, our focus is on the challenges posed by OOD data in real-world applications and on enhancing the robustness and generalization of models. In this paper, we propose an infrared-visible fusion framework based on Multi-View Augmentation. For external data augmentation, Top-k Selective Vision Alignment is employed to mitigate distribution shifts between datasets by performing RGB-wise transformations on visible images. This strategy effectively introduces augmented samples, enhancing the adaptability of the model to complex real-world scenarios. Additionally, for internal data augmentation, self-supervised learning is established using Weak-Aggressive Augmentation. This enables the model to learn more robust and general feature representations during the fusion process, thereby improving robustness and generalization. Extensive experiments demonstrate that the proposed method exhibits superior performance and robustness across various conditions and environments. Our approach significantly enhances the reliability and stability of IVIF tasks in practical applications.

图像融合多模态鲁棒性自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。