用遮罩提示实现红外可见光图像可控融合,提升智能系统感知能力
CtrlFuse: Mask-Prompt Guided Controllable Infrared and Visible Image Fusion
- 通过遮罩提示动态引导融合过程,实现交互式控制
- 在融合可控性和分割准确率上均达当前最佳水平
- 适合需要灵活感知特定目标的无人机等无人系统
红外与可见光图像融合通过结合互补模态,生成全天候感知能力的图像,增强智能无人系统的环境认知。现有方法或仅关注像素级融合而忽视下游任务适应性,或通过级联检测/分割模型隐式学习固定语义,难以应对多样化的语义目标感知需求。我们提出CtrlFuse,一种可交互控制的图像融合框架,基于遮罩提示实现动态融合。模型包含多模态特征提取器、参考提示编码器(RPE)和提示-语义融合模块(PSFM)。RPE通过输入遮罩引导微调预训练分割模型,动态编码任务相关语义;PSFM则显式将这些语义注入融合特征。通过并行分割与融合分支的协同优化,任务性能与融合质量相互提升。实验表明,该方法在融合可控性和分割精度上均达到领先水平,适配后的任务分支甚至超过原始分割模型。
原文摘要 · Abstract (English)
Infrared and visible image fusion generates all-weather perception-capable images by combining complementary modalities, enhancing environmental awareness for intelligent unmanned systems. Existing methods either focus on pixel-level fusion while overlooking downstream task adaptability or implicitly learn rigid semantics through cascaded detection/segmentation models, unable to interactively address diverse semantic target perception needs. We propose CtrlFuse, a controllable image fusion framework that enables interactive dynamic fusion guided by mask prompts. The model integrates a multi-modal feature extractor, a reference prompt encoder (RPE), and a prompt-semantic fusion module (PSFM). The RPE dynamically encodes task-specific semantic prompts by fine-tuning pre-trained segmentation models with input mask guidance, while the PSFM explicitly injects these semantics into fusion features. Through synergistic optimization of parallel segmentation and fusion branches, our method achieves mutual enhancement between task performance and fusion quality. Experiments demonstrate state-of-the-art results in both fusion controllability and segmentation accuracy, with the adapted task branch even outperforming the original segmentation model.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。