arXiv:2510.02599cs.CV2025-10

无需训练,优化提示词嵌入即可提升图像美感。

PEO: Training-Free Aesthetic Quality Enhancement in Pre-Trained Text-to-Image Diffusion Models with Prompt Embedding Optimization

  • 通过优化提示词嵌入,提升生成图像的审美质量。
  • 在不改变原始提示的前提下,显著改善图像视觉效果。
  • 无需微调模型,适用于任意预训练文生图模型。

本文提出一种新方法,针对给定简单提示,在预训练文生图扩散模型中实现美学质量提升。该方法称为提示词嵌入优化(PEO),以预训练文生图扩散模型为骨干,通过优化原始提示的文本嵌入来增强生成图像的视觉质量。其采用三重目标函数:提升生成图像的美学保真度、保证与优化后文本嵌入的一致性,以及最小化与初始提示的偏离。后者通过提示保留项实现。PEO具有无需训练、对骨干模型无依赖的特点。定量与定性评估均表明,该方法性能超越或媲美当前最优的文生图及提示适配方法。

原文摘要 · Abstract (English)

This paper introduces a novel approach to aesthetic quality improvement in pre-trained text-to-image diffusion models when given a simple prompt. Our method, dubbed Prompt Embedding Optimization (PEO), leverages a pre-trained text-to-image diffusion model as a backbone and optimizes the text embedding of a given simple and uncurated prompt to enhance the visual quality of the generated image. We achieve this by a tripartite objective function that improves the aesthetic fidelity of the generated image, ensures adherence to the optimized text embedding, and minimal divergence from the initial prompt. The latter is accomplished through a prompt preservation term. Additionally, PEO is training-free and backbone-independent. Quantitative and qualitative evaluations confirm the effectiveness of the proposed method, exceeding or equating the performance of state-of-the-art text-to-image and prompt adaptation methods.

文生图美学增强提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。