用扩散模型一键生成图文一致的高质量设计图
DesignDiffusion: High-Quality Text-to-Design Image Generation with Diffusion Models
- 端到端扩散模型直接合成文本与视觉元素
- 字符嵌入和定位损失提升文字生成准确率
- 自对弈偏好优化使图文风格更统一
本文提出DesignDiffusion,一个用于从文本描述生成设计图像的新框架。主要挑战在于保持文本与视觉内容的准确性和风格一致性。现有视觉文本生成方法通常局限于特定区域生成文字,限制了创意表达,导致在设计图像生成中出现图文风格或色彩不一致的问题。为此,我们提出一种无需复杂布局建模的端到端单阶段扩散框架,直接根据用户提示合成文本与视觉元素。该框架利用基于视觉文本的独特色符嵌入增强输入提示,并引入字符定位损失以加强文本生成过程中的监督。此外,采用自对弈直接偏好优化(self-play DPO)微调策略,进一步提升生成图文的质量与准确性。大量实验表明,DesignDiffusion在设计图像生成任务上达到当前最优性能。
原文摘要 · Abstract (English)
In this paper, we present DesignDiffusion, a simple yet effective framework for the novel task of synthesizing design images from textual descriptions. A primary challenge lies in generating accurate and style-consistent textual and visual content. Existing works in a related task of visual text generation often focus on generating text within given specific regions, which limits the creativity of generation models, resulting in style or color inconsistencies between textual and visual elements if applied to design image generation. To address this issue, we propose an end-to-end, one-stage diffusion-based framework that avoids intricate components like position and layout modeling. Specifically, the proposed framework directly synthesizes textual and visual design elements from user prompts. It utilizes a distinctive character embedding derived from the visual text to enhance the input prompt, along with a character localization loss for enhanced supervision during text generation. Furthermore, we employ a self-play Direct Preference Optimization fine-tuning strategy to improve the quality and accuracy of the synthesized visual text. Extensive experiments demonstrate that DesignDiffusion achieves state-of-the-art performance in design image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。