arXiv:2501.07070cs.CV2025-01中稿 · ICASSP 2025, Githu…被引 10

通过分阶段提示提升扩散模型图像生成的精准度。

Enhancing Image Generation Fidelity via Progressive Prompts

  • 用大语言模型生成图像的高层与低层描述,指导分阶段生成。
  • 发现深层交叉注意力控制高层内容,浅层控制细节。
  • 支持区域提示控制,适合需要精细调控的生成任务。

扩散变换器(DiT)架构在图像生成中备受关注,实现了更高的保真度、性能和多样性。然而,现有基于DiT的方法多聚焦全局感知合成,对局部提示控制的研究较少。本文提出一种从粗到细的区域提示跟随生成流程:首先利用大语言模型(LLM)生成图像的高层(如内容、主题、物体)和低层(如细节、风格)描述;接着探究不同深度交叉注意力层的影响,发现深层始终负责高层内容控制,浅层则处理低层内容控制。将多种提示注入所提出的区域交叉注意力控制模块,实现从粗到细的生成。实验表明,该方法显著提升了基于DiT的图像生成可控性,定量与定性结果均验证了其有效性。

原文摘要 · Abstract (English)

The diffusion transformer (DiT) architecture has attracted significant attention in image generation, achieving better fidelity, performance, and diversity. However, most existing DiT - based image generation methods focus on global - aware synthesis, and regional prompt control has been less explored. In this paper, we propose a coarse - to - fine generation pipeline for regional prompt - following generation. Specifically, we first utilize the powerful large language model (LLM) to generate both high - level descriptions of the image (such as content, topic, and objects) and low - level descriptions (such as details and style). Then, we explore the influence of cross - attention layers at different depths. We find that deeper layers are always responsible for high - level content control, while shallow layers handle low - level content control. Various prompts are injected into the proposed regional cross - attention control for coarse - to - fine generation. By using the proposed pipeline, we enhance the controllability of DiT - based image generation. Extensive quantitative and qualitative results show that our pipeline can improve the performance of the generated images.

图像生成扩散模型提示控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。