arXiv:2503.15283cs.CV2025-03ICCV被引 6

无需训练即可实现图文到图像生成,提升复杂指令下的生成质量

TF-TI2I: Training-Free Text-and-Image-to-Image Generation via Multi-Modal Implicit-Context Learning in Text-to-Image Models

  • 利用文本与图像隐式上下文交互,通过压缩视觉表征选择性融合信息
  • 在多个基准上优于现有方法,复杂指令下生成质量稳定
  • 适用于需要快速部署的图像编辑场景,尤其适合多参考输入任务

图文到图像(TI2I)是文本到图像(T2I)的扩展,通过结合图像输入与文本指令增强图像生成能力。现有方法常仅部分利用图像输入,关注特定元素如对象或风格,或在复杂多图指令下生成质量下降。为此,我们提出无需训练的图文到图像生成方法(TF-TI2I),适配SD3等先进T2I模型。该方法基于多模态迪特(MM-DiT)架构,发现文本标记可隐式学习视觉标记中的信息。通过从参考图像中提取紧凑视觉表示,并采用参考上下文掩码技术,仅允许与指令相关的视觉信息参与上下文交互;同时引入胜者为王模块,通过优先选择最相关参考来缓解分布偏移。为填补评估空白,我们还构建了兼容现有T2I方法的FG-TI2I Bench基准。实验表明,该方法在多种基准上表现稳健,有效应对复杂生成任务。

原文摘要 · Abstract (English)

Text-and-Image-To-Image (TI2I), an extension of Text-To-Image (T2I), integrates image inputs with textual instructions to enhance image generation. Existing methods often partially utilize image inputs, focusing on specific elements like objects or styles, or they experience a decline in generation quality with complex, multi-image instructions. To overcome these challenges, we introduce Training-Free Text-and-Image-to-Image (TF-TI2I), which adapts cutting-edge T2I models such as SD3 without the need for additional training. Our method capitalizes on the MM-DiT architecture, in which we point out that textual tokens can implicitly learn visual information from vision tokens. We enhance this interaction by extracting a condensed visual representation from reference images, facilitating selective information sharing through Reference Contextual Masking -- this technique confines the usage of contextual tokens to instruction-relevant visual information. Additionally, our Winner-Takes-All module mitigates distribution shifts by prioritizing the most pertinent references for each vision token. Addressing the gap in TI2I evaluation, we also introduce the FG-TI2I Bench, a comprehensive benchmark tailored for TI2I and compatible with existing T2I methods. Our approach shows robust performance across various benchmarks, confirming its effectiveness in handling complex image-generation tasks.

图文生成无训练图像编辑多模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。