用多粒度文本指导图像融合,提升光照与对焦差异下的画质。
Multi-Grained Text-Guided Image Fusion for Multi-Exposure and Multi-Focus Scenarios
- 分层文本描述细粒度、结构和语义信息,引导跨模态融合
- 多粒度监督增强视觉与文本特征对齐,提升融合精度
- 通过显著性增强数据,强化细节保留,适合图像增强场景
图像融合旨在从不同曝光或对焦条件下拍摄的输入图像中合成一张高质量图像。核心挑战在于有效处理输入间动态范围和聚焦深度的差异。随着视觉-语言模型的发展,近期方法引入文本描述作为辅助指导以提升融合质量。然而,简单使用粗粒度描述会削弱对细粒度细节的理解,并导致跨模态对齐困难。为此,我们提出多粒度文本引导图像融合(MTIF),包含三项关键设计:首先,引入多粒度文本描述,分别捕捉细粒度细节、结构线索和语义内容,通过层次化跨模态调制模块指导融合;其次,在每个粒度层级引入监督信号,促进视觉与文本特征对齐,增强辅助文本的实用性;第三,采用显著性驱动的数据增强模块,以密集语义内容扩充训练数据,进一步强化跨模态调制与对齐。大量实验表明,MTIF在多曝光和多对焦图像融合任务上均持续优于现有方法。
原文摘要 · Abstract (English)
Image fusion aims to synthesize a single high-quality image from a pair of inputs captured under challenging conditions, such as differing exposure levels or focal depths. A core challenge lies in effectively handling disparities in dynamic range and focus depth between the inputs. With the advent of vision-language models, recent methods incorporate textual descriptions as auxiliary guidance to enhance fusion quality. However, simply incorporating coarse-grained descriptions hampers the understanding of fine-grained details and poses challenges for precise cross-modal alignment. To address these limitations, we propose Multi-grained Text-guided Image Fusion (MTIF), a novel fusion paradigm with three key designs. First, it introduces multi-grained textual descriptions that separately capture fine details, structural cues, and semantic content, guiding image fusion through a hierarchical cross-modal modulation module. Second, it involves supervision signals at each granularity to facilitate alignment between visual and textual features and enhance the utility of auxiliary text. Third, it adopts a saliency-driven enrichment module to augment training data with dense semantic content, further strengthening the cross-modal modulation and alignment. Extensive experiments show that MTIF consistently outperforms previous methods on both multi-exposure and multi-focus image fusion tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。