用轻量模型自动优化文本生成图像的提示词,提升画面质量与一致性。
TIPO: Text to Image with Text Presampling for Prompt Optimization
- 基于轻量预训练模型从原始提示词生成更丰富的描述。
- 在多领域实验中显著降低视觉瑕疵,人类偏好率更高。
- 无需大模型或强化学习,计算高效适合大规模应用。
TIPO(文本到图像提示词优化)提出一种高效的自动提示词优化方法,用于文本到图像生成任务。从简单的用户输入出发,TIPO利用一个轻量级预训练模型将提示词扩展为更丰富、更详细的版本。其核心思想是从未知语义空间中采样特定子分布的优化提示词,在保留原意的同时显著提升图像质量、连贯性和细节表现。相比依赖大型语言模型或强化学习的高资源方法,TIPO具有更强的计算效率和可扩展性,为自动化提示工程提供了新路径。在多个领域的大量实验表明,TIPO实现了更强的文本对齐能力、更低的视觉伪影,并在人类偏好测试中持续领先,同时保持了优异的美学质量。结果验证了分布对齐提示工程的有效性,并指向文本到图像生成中更广泛、可扩展的自动化优化前景。
原文摘要 · Abstract (English)
TIPO (Text-to-Image Prompt Optimization) introduces an efficient approach for automatic prompt refinement in text-to-image (T2I) generation. Starting from simple user prompts, TIPO leverages a lightweight pre-trained model to expand these prompts into richer and more detailed versions. Conceptually, TIPO samples refined prompts from a targeted sub-distribution within the broader semantic space, preserving the original intent while significantly improving visual quality, coherence, and detail. Unlike resource-intensive methods based on large language models (LLMs) or reinforcement learning (RL), TIPO offers strong computational efficiency and scalability, opening new possibilities for effective automated prompt engineering in T2I tasks. Extensive experiments across multiple domains demonstrate that TIPO achieves stronger text alignment, reduced visual artifacts, and consistently higher human preference rates, while maintaining competitive aesthetic quality. These results highlight the effectiveness of distribution-aligned prompt engineering and point toward broader opportunities for scalable, automated refinement in text-to-image generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。