用文本指导图像融合,应对复杂退化下的信息丢失问题。
Unleashing the Power of Text: Text-Guided Flow Matching for Image Fusion under Complex Degradations

- 将文本描述转为结构化提示,动态引导跨模态融合
- 在多个基准上优于或媲美现有方法,保持细节与热源特征
- 适合处理真实场景中多种退化叠加的红外可见光图像融合
在真实退化场景下进行红外-可见光图像融合是一项挑战,因为退化不仅导致观测图像中可靠的模态特异性信息丢失,还阻碍融合过程。最新研究表明,文本可提供关于退化特性的先验信息,弥补受损输入图像中有限的证据,促进融合。然而,现有方法通常将固定的全局文本表示注入视觉特征,难以适应空间变化的退化、局部结构和热显著性。为此,我们提出 TGFusion,一种文本引导的潜在空间流匹配框架,统一退化抑制与跨模态融合。TGFusion 将任务、退化和生成线索编码为结构化提示。为充分挖掘这些先验,设计了提示条件的多流联合流变换器,将文本作为独立语义流,与融合、可见光和红外流并行。联合注意力实现语义与视觉表示间的令牌级双向交互和逐层更新,使退化语义能动态引导可靠信息选择与融合潜在表示生成。在多个公开基准和复杂退化场景下的大量实验表明,TGFusion 在感知质量、图像自然度、结构细节保留和红外显著性保持方面均达到优越或具有竞争力的性能,且对单个和复合退化均保持鲁棒性。
原文摘要 · Abstract (English)
Infrared-visible image fusion under realistic degradation scenarios is a challenging task, as degradations not only cause a loss of reliable modality-specific information in observed images but also hinder the fusion process. Recent studies indicate that text can provide prior information about degradation characteristics, complementing the limited evidence available from corrupted input images and facilitating fusion. However, existing methods typically inject fixed global text representations into visual features, making it difficult for textual guidance to adapt to spatially varying degradations, local structures, and thermal saliency. To this end, we propose TGFusion, a text-guided latent-space flow matching framework that unifies degradation suppression and cross-modal fusion. TGFusion encodes task, degradation, and generation cues into structured prompts. To fully exploit these priors, we design a Prompt-conditioned Multi-stream Joint Flow Transformer that represents text as an independent semantic stream alongside fusion, visible, and infrared streams. Joint attention enables token-level bidirectional interaction and layer-wise updating among semantic and visual representations, allowing degradation semantics to dynamically guide reliable information selection and fusion latent generation. Extensive experiments on public benchmarks and complex degradation scenarios demonstrate that TGFusion achieves superior or competitive performance in perceptual quality, image naturalness, structural-detail preservation, and infrared-saliency retention, while remaining robust across diverse single and compound degradations.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。