用语言指令控制图像融合,实现多模态信息精准整合。
Towards Unified Semantic and Controllable Image Fusion: A Diffusion Transformer Approach
- 通过扩散变换器联合编码图像与指令,实现语义感知融合
- 无需真实融合图像即可训练,支持多场景零样本泛化
- 统一处理红外可见光、多焦点等融合任务,支持用户分级控制
图像融合旨在融合多传感器模态的互补信息,但现有方法在鲁棒性、适应性和可控性方面仍受限。多数网络针对特定任务设计,难以灵活融入用户意图,尤其在低光照退化、色彩偏移或曝光不平衡等复杂场景下表现不佳。此外,缺乏真实融合图像和现有数据集规模小,使得端到端模型同时理解高层语义与进行细粒度多模态对齐变得困难。为此,我们提出DiTFuse——一种由指令驱动的扩散变换器(DiT)框架,可在单一模型中实现端到端、语义感知的图像融合。通过将两张图像与自然语言指令共同嵌入共享潜在空间,DiTFuse实现了层级化、细粒度的融合控制,突破了传统预融合与后融合管道难以注入高层语义的瓶颈。训练阶段采用多退化掩码图像建模策略,使网络联合学习跨模态对齐、模态不变修复与任务感知特征选择,无需依赖真实融合图像。一个精心构建的多粒度指令数据集进一步赋予模型交互式融合能力。DiTFuse统一处理红外-可见光、多焦点、多曝光融合,以及文本控制的精细化调整和下游任务。在公开的IVIF、MFF、MEF基准上的实验验证其在定量与定性指标上均优于现有方法,纹理更清晰,语义保留更好。模型还支持多级用户控制,并具备向其他多图像融合场景的零样本泛化能力,包括指令条件分割。
原文摘要 · Abstract (English)
Image fusion aims to blend complementary information from multiple sensing modalities, yet existing approaches remain limited in robustness, adaptability, and controllability. Most current fusion networks are tailored to specific tasks and lack the ability to flexibly incorporate user intent, especially in complex scenarios involving low-light degradation, color shifts, or exposure imbalance. Moreover, the absence of ground-truth fused images and the small scale of existing datasets make it difficult to train an end-to-end model that simultaneously understands high-level semantics and performs fine-grained multimodal alignment. We therefore present DiTFuse, instruction-driven Diffusion-Transformer (DiT) framework that performs end-to-end, semantics-aware fusion within a single model. By jointly encoding two images and natural-language instructions in a shared latent space, DiTFuse enables hierarchical and fine-grained control over fusion dynamics, overcoming the limitations of pre-fusion and post-fusion pipelines that struggle to inject high-level semantics. The training phase employs a multi-degradation masked-image modeling strategy, so the network jointly learns cross-modal alignment, modality-invariant restoration, and task-aware feature selection without relying on ground truth images. A curated, multi-granularity instruction dataset further equips the model with interactive fusion capabilities. DiTFuse unifies infrared-visible, multi-focus, and multi-exposure fusion-as well as text-controlled refinement and downstream tasks-within a single architecture. Experiments on public IVIF, MFF, and MEF benchmarks confirm superior quantitative and qualitative performance, sharper textures, and better semantic retention. The model also supports multi-level user control and zero-shot generalization to other multi-image fusion scenarios, including instruction-conditioned segmentation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。