arXiv:2605.01916cs.CV2026-05被引 1

用自演化先验动态生成卷积核,提升红外可见光图像融合效果

EAPFusion: Intrinsic Evolving Auxiliary Prior Guidance for Infrared and Visible Image Fusion

论文配图:EAPFusion: Intrinsic Evolving Auxiliary Prior Guidance for Infrared and Visible Image Fusion
图 1 · 摘自论文原文
  • 通过自演化内在先验动态生成卷积核,实现实例自适应
  • 在多个数据集上达到顶尖融合效果,显著提升下游语义分割性能
  • 适合需要高精度跨模态融合的夜视感知场景

红外-可见光图像融合旨在结合红外热敏特征与可见光细节纹理,生成信息丰富的融合图像。该技术对复杂场景下的真实感知任务(如夜间自动驾驶、搜救、监控)至关重要,并可进一步提升语义分割等下游任务表现。然而,现有方法依赖静态训练权重,无法在推理时适应特定场景内容,且注入粗粒度辅助语义时常导致粒度不匹配,难以同时突出目标并保留细节。为此,本文提出EAPFusion,采用自演化内在先验替代外部辅助模型。具体而言,EAPFusion维护一组紧凑的内在先验,并在多尺度间逐步更新;这些演化先验用于动态生成卷积核,将固定预训练滤波器范式转变为先验条件驱动的实例自适应参数。此外,设计通道级融合模块,对红外与可见光通道进行局部混洗与交错,增强跨模态互补性。在多个数据集(包括跨数据集评估)上的实验表明,该方法在定量与定性指标上均达到当前最优,且持续提升下游任务性能。

原文摘要 · Abstract (English)

Infrared-visible image fusion aims to create an information-rich fused image by integrating the complementary thermal saliency from infrared sensing and fine textures from visible imaging. Such accurate fusion is essential for real-world perception applications in complex scenes, including nighttime autonomous driving, search and rescue, and surveillance, and can further benefit downstream tasks such as semantic segmentation. However, most existing fusion methods rely upon static trained weights that cannot adapt to scene-specific content at inference time, and often suffer from a granularity mismatch when coarse auxiliary semantics are injected, which makes it difficult to simultaneously highlight targets and preserve details. In this work, we propose EAPFusion to address these issues by using self-evolving intrinsic priors instead of relying on external auxiliary models. Concretely, EAPFusion maintains a compact set of intrinsic priors and progressively updates them across scales. These evolved priors are utilized to dynamically generate convolutional kernels, shifting the paradigm from fixed, pre-trained filters to instance-adaptive parameters via prior-conditioned dynamic convolution. Furthermore, we design a channel-level fusion module that shuffles and interleaves infrared and visible channels, applying local channel mixing to boost cross-modal complementarity. Experiments on different datasets, including cross-dataset evaluation and semantic segmentation, show that the proposed method achieves state-of-the-art quantitative and qualitative fusion results, and consistently boosts downstream performance. Code is coming soon.

图像融合跨模态自适应卷积红外可见光

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。