用低层视觉任务提升多模态图像融合,单模型搞定多种任务
One Model for ALL: Low-Level Task Interaction Is a Key to Task-Agnostic Image Fusion
- 基于像素级监督的低层特征交互,避免语义鸿沟
- 单一模型在已见与未见场景中均表现优异
- 不仅支持融合,还能增强单模态图像
现有高级图像融合方法多关注高层任务,任务间交互受语义差异影响,需复杂桥梁机制。本文提出利用数字摄影中的低层视觉任务,通过像素级监督实现有效特征交互。该新范式为无监督多模态融合提供强指导,无需依赖抽象语义,增强任务共享特征学习,提升通用性。得益于混合图像特征和强化的通用表征,所提GIFNet支持多样融合任务,在已见与未见场景中均取得高性能。实验表明,该框架还可用于单模态图像增强,展现更强实用性。代码将公开于https://github.com/AWCXV/GIFNet。
原文摘要 · Abstract (English)
Advanced image fusion methods mostly prioritise high-level missions, where task interaction struggles with semantic gaps, requiring complex bridging mechanisms. In contrast, we propose to leverage low-level vision tasks from digital photography fusion, allowing for effective feature interaction through pixel-level supervision. This new paradigm provides strong guidance for unsupervised multimodal fusion without relying on abstract semantics, enhancing task-shared feature learning for broader applicability. Owning to the hybrid image features and enhanced universal representations, the proposed GIFNet supports diverse fusion tasks, achieving high performance across both seen and unseen scenarios with a single model. Uniquely, experimental results reveal that our framework also supports single-modality enhancement, offering superior flexibility for practical applications. Our code will be available at https://github.com/AWCXV/GIFNet.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。