arXiv:2509.26641cs.CV2025-09被引 14

将视觉语言模型与扩散模型分离,用语义上下文提升图像生成与编辑的准确性和一致性。

Query-Kontext: An Unified Multimodal Model for Image Generation and Editing

  • 通过多模态'上下文'令牌连接VLM与扩散模型,分离理解与生成任务。
  • 三阶段训练使模型在图像生成和指令编辑上达到或超越现有顶尖方法。
  • 适合需要高保真图像生成与精准指令响应的研究者和开发者。

统一多模态模型(UMMs)在文本到图像生成(T2I)和编辑(TI2I)任务中表现卓越,现有框架或采用视觉语言模型(VLM)与扩散生成器耦合,或早期融合理解与生成模态。本文指出,当前统一框架中,多模态生成推理能力(包括指令理解、定位与身份保持)与高保真合成被内在纠缠。为此,我们提出Query-Kontext,通过由语义线索和粗粒度图像条件组成的多模态'上下文',连接强大的VLM与扩散模型。该设计将复杂生成推理交给VLM,保留扩散模型专注于高质量视觉合成。我们提出三阶段渐进式训练策略:首先,用轻量级扩散头连接VLM与多模态上下文令牌以释放其生成推理能力;其次,扩展为大型预训练扩散模型以增强细节与真实感;最后,引入低层图像编码器提升图像保真度,并在下游任务上进行指令微调。此外,我们构建了涵盖真实、合成与开源数据集的综合数据管道,覆盖图像生成、指令驱动编辑、定制化生成及多主体组合等多样场景。实验表明,本方法在多个任务上匹配强基线,甚至优于特定任务的最先进模型。

原文摘要 · Abstract (English)

Unified Multimodal Models (UMMs) have demonstrated remarkable performance in text-to-image generation (T2I) and editing (TI2I), whether instantiated as assembled unified frameworks which couple powerful vision-language model (VLM) with diffusion-based generator, or as naive Unified Multimodal Models with an early fusion of understanding and generation modalities. We contend that in current unified frameworks, the crucial capability of multimodal generative reasoning which encompasses instruction understanding, grounding, and image referring for identity preservation and faithful reconstruction, is intrinsically entangled with high-fidelity synthesis. In this work, we introduce Query-Kontext, a novel approach that bridges the VLM and diffusion model via a multimodal ``kontext'' composed of semantic cues and coarse-grained image conditions encoded from multimodal inputs. This design delegates the complex ability of multimodal generative reasoning to powerful VLM while reserving diffusion model's role for high-quality visual synthesis. To achieve this, we propose a three-stage progressive training strategy. First, we connect the VLM to a lightweight diffusion head via multimodal kontext tokens to unleash the VLM's generative reasoning ability. Second, we scale this head to a large, pre-trained diffusion model to enhance visual detail and realism. Finally, we introduce a low-level image encoder to improve image fidelity and perform instruction tuning on downstream tasks. Furthermore, we build a comprehensive data pipeline integrating real, synthetic, and open-source datasets, covering diverse multimodal reference-to-image scenarios, including image generation, instruction-driven editing, customized generation, and multi-subject composition. Experiments show that our approach matches strong unified baselines and even outperforms task-specific state-of-the-art methods in several cases.

图像生成多模态扩散模型指令编辑

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。