arXiv:2412.18928cs.CVcs.LG2024-12CVPR被引 10

一个模型搞定多种图像控制,用图文指令统一生成

UNIC-Adapter: Unified Image-instruction Adapter with Multi-modal Transformer for Image Generation

  • 用多模态变换器融合条件图与文本指令,统一控制生成
  • 在空间布局、主体生成和风格迁移任务中表现优异
  • 适合需要灵活图像编辑的AI设计师或开发者

近年来,文本到图像生成模型取得显著进展,尤其扩散模型能基于文本描述生成高质量图像。然而,仅依赖文本提示时,模型在像素级布局、物体外观和全局风格控制上仍存在不足。为解决此问题,已有工作引入条件图像作为辅助输入以增强控制,但通常需为不同参考输入设计专用模型。本文提出一种统一图像指令适配器(UNIC-Adapter),基于多模态扩散变换器架构,实现单一框架内对多种条件的灵活可控生成。UNIC-Adapter通过融合条件图像与任务指令,利用改进的旋转位置编码交叉注意力机制注入多模态指令信息。在像素级空间控制、主体驱动生成及基于风格图像的合成等任务上的实验表明,该方法在统一可控图像生成方面具有显著效果。

原文摘要 · Abstract (English)

Recently, text-to-image generation models have achieved remarkable advancements, particularly with diffusion models facilitating high-quality image synthesis from textual descriptions. However, these models often struggle with achieving precise control over pixel-level layouts, object appearances, and global styles when using text prompts alone. To mitigate this issue, previous works introduce conditional images as auxiliary inputs for image generation, enhancing control but typically necessitating specialized models tailored to different types of reference inputs. In this paper, we explore a new approach to unify controllable generation within a single framework. Specifically, we propose the unified image-instruction adapter (UNIC-Adapter) built on the Multi-Modal-Diffusion Transformer architecture, to enable flexible and controllable generation across diverse conditions without the need for multiple specialized models. Our UNIC-Adapter effectively extracts multi-modal instruction information by incorporating both conditional images and task instructions, injecting this information into the image generation process through a cross-attention mechanism enhanced by Rotary Position Embedding. Experimental results across a variety of tasks, including pixel-level spatial control, subject-driven image generation, and style-image-based image synthesis, demonstrate the effectiveness of our UNIC-Adapter in unified controllable image generation.

图像生成多模态扩散模型可控生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。