让红外可见光融合图像按任务需求自动调整,无需重新训练。
Instruction-Driven Fusion of Infrared-Visible Images: Tailoring for Diverse Downstream Tasks
- 根据用户指令生成动态提示,引导特征提取适配具体任务。
- 在目标检测、语义分割等任务上表现优异,性能优于传统方法。
- 适合需要多任务兼容的实时应用,如智能监控与自动驾驶。
红外与可见光图像融合技术的核心价值在于支持下游任务。然而,现有方法在处理多个下游任务时面临训练复杂度高、单任务性能显著下降的问题。为此,我们提出面向任务的自适应调节(T-OAR)机制,并引入任务相关动态提示注入(T-DPI)模块,该模块基于用户输入的文本指令生成任务特定的动态提示,融入目标表征中,引导特征提取模块生成更贴合任务需求的表示。将T-DPI集成到T-OAR框架后,所提方法能在不进行单独训练或使用任务专用权重的情况下,生成针对不同任务定制的融合图像。这不仅降低计算开销,还提升多任务环境下的适应性与性能。实验表明,该方法在目标检测、语义分割和显著性物体检测任务中均表现突出,展现出强大的适应性、灵活性与任务特异性,为多任务场景下的图像融合提供了高效解决方案,具有广泛的应用潜力。
原文摘要 · Abstract (English)
The primary value of infrared and visible image fusion technology lies in applying the fusion results to downstream tasks. However, existing methods face challenges such as increased training complexity and significantly compromised performance of individual tasks when addressing multiple downstream tasks simultaneously. To tackle this, we propose Task-Oriented Adaptive Regulation (T-OAR), an adaptive mechanism specifically designed for multi-task environments. Additionally, we introduce the Task-related Dynamic Prompt Injection (T-DPI) module, which generates task-specific dynamic prompts from user-input text instructions and integrates them into target representations. This guides the feature extraction module to produce representations that are more closely aligned with the specific requirements of downstream tasks. By incorporating the T-DPI module into the T-OAR framework, our approach generates fusion images tailored to task-specific requirements without the need for separate training or task-specific weights. This not only reduces computational costs but also enhances adaptability and performance across multiple tasks. Experimental results show that our method excels in object detection, semantic segmentation, and salient object detection, demonstrating its strong adaptability, flexibility, and task specificity. This provides an efficient solution for image fusion in multi-task environments, highlighting the technology's potential across diverse applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。