仅用一张模糊可见光图,生成媲美多模态融合的跨谱场景表示。
MagicFuse: Single Image Fusion for Visual and Semantic Reinforcement
- 基于扩散模型构建知识增强与跨谱生成双分支。
- 从单张低质可见光图中还原热辐射分布并生成跨谱表征。
- 兼顾视觉质量与语义准确性,适合复杂环境感知任务。
本文聚焦于仅有可见光传感器时如何在恶劣条件下仍获得多模态图像融合的优势。为此,提出单图融合新范式,将传统数据级融合拓展至知识级。我们设计MagicFuse框架,仅需一张低质量可见光图像即可生成完整的跨谱场景表征。该框架包含两个核心分支:基于扩散模型的同谱知识强化分支,挖掘可见光中被遮蔽的场景信息;以及跨谱知识生成分支,学习热辐射分布向红外谱的映射规律。进一步,设计多域知识融合分支,融合两分支扩散流的随机噪声,通过逐步采样获得跨谱表征。同时引入视觉与语义双重约束,确保结果既符合人眼观察又支持下游语义决策。大量实验表明,即使仅依赖单张劣质可见光图像,MagicFuse在视觉与语义表现上仍可媲美甚至超越采用多模态输入的先进融合方法。代码已开源:https://github.com/zhayanping/MagicFuse。
原文摘要 · Abstract (English)
This paper focuses on a highly practical scenario: how to continue benefiting from the advantages of multi-modal image fusion under harsh conditions when only visible imaging sensors are available. To achieve this goal, we propose a novel concept of single-image fusion, which extends conventional data-level fusion to the knowledge level. Specifically, we develop MagicFuse, a novel single image fusion framework capable of deriving a comprehensive cross-spectral scene representation from a single low-quality visible image. MagicFuse first introduces an intra-spectral knowledge reinforcement branch and a cross-spectral knowledge generation branch based on the diffusion models. They mine scene information obscured in the visible spectrum and learn thermal radiation distribution patterns transferred to the infrared spectrum, respectively. Building on them, we design a multi-domain knowledge fusion branch that integrates the probabilistic noise from the diffusion streams of these two branches, from which a cross-spectral scene representation can be obtained through successive sampling. Then, we impose both visual and semantic constraints to ensure that this scene representation can satisfy human observation while supporting downstream semantic decision-making. Extensive experiments show that our MagicFuse achieves visual and semantic representation performance comparable to or even better than state-of-the-art fusion methods with multi-modal inputs, despite relying solely on a single degraded visible image. The code is publicly available at https://github.com/zhayanping/MagicFuse.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。