arXiv:2608.13045cs.CV2026-08中稿 · ECCV

用动态提示引导红外可见光图像融合,提升画质与检测效果

P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

论文配图:P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation
图 1 · 摘自论文原文
  • 通过双内在提示蒸馏,让模型自适应调节模态竞争
  • 在5个数据集上14项指标达顶尖水平,显著提升融合质量
  • 适合做多模态感知、目标检测增强的工程师和研究者

红外-可见光图像融合(IVIF)对多模态感知至关重要,但热成像与纹理特征之间的固有差异仍是核心挑战。现有基于先验的方法多依赖静态约束,导致优化冲突;或使用外部语义先验(如CLIP/DINO),难以捕捉模态本质特性。为此,我们提出P2Fusion,一种基于双内在提示蒸馏的融合框架。不采用硬编码惩罚,而是将热敏感性与空间质量等图像内在先验蒸馏为可学习的动态调控器。通过教-融机制实现双粒度渐进引导,并引入门控动态专家重校准(GDER)模块,解耦特征优化,促进专家专业化。大量实验表明,P2Fusion在五个主流数据集上达到领先性能,14/20关键指标优于现有方法。同时显著提升下游感知鲁棒性:在MSRS上目标检测mAP提升+3.2%,M3FD上+0.5%,DroneVehicle上+0.9%。

原文摘要 · Abstract (English)

Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.g., CLIP/DINO), which frequently fail to exploit the intrinsic modality characteristics essential for high-fidelity fusion. To address these issues, we propose P2Fusion, a prior-guided distillation-based framework that reformulates IVIF via dual intrinsic prompts. Instead of imposing hard-coded penalties, we distill image-intrinsic priors, thermal saliency and spatial quality, into learnable dynamic regulators. Specifically, a Teach-to-Fuse mechanism provides dual-granularity progressive guidance, coupled with a Gated Dynamic Expert Recalibration (GDER) module for decoupled feature refinement. This design enables the network to adaptively mediate modal competition through expert specialization. Extensive experiments demonstrate that P2Fusion achieves state-of-the-art performance across five mainstream datasets. Notably, our framework demonstrates consistent performance advantages in fusion quality, achieving state-of-the-art results in 14 out of 20 key evaluation metrics across 5 benchmarks. Furthermore, it effectively contributes to the robustness of downstream perception, such as +3.2% mAP on MSRS, +0.5% mAP on M3FD and +0.9% mAP on DroneVehicle for object detection. Our code will be available at https://github.com/YiShi99/P2Fusion

图像融合多模态提示学习目标检测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。