arXiv:2603.12506cs.CVcs.AI2026-03

用提示词评估提升文生图质量,减少试错次数。

Naïve PAINE: Lightweight Text-to-Image Generation Improvement with Prompt Evaluation

  • 从初始噪声和提示词直接预测图像质量
  • 筛选高质量噪声,显著提升生成结果满意度
  • 轻量级设计,可无缝接入现有文生图流程

文生图生成主要依赖扩散模型(DM),其基于随机高斯噪声,相同输入会产生不同结果,用户需多次尝试才能获得满意图像,造成‘赌徒负担’。尽管扩散模型使用随机采样,但生成内容质量高度依赖提示词与模型对提示的生成能力。为此,我们提出轻量级方法 Naïve PAINE,利用文生图偏好基准直接预测图像质量,从初始噪声中筛选出高质量候选,再交由扩散模型生成。该方法还提供针对提示词的生成质量反馈。实验表明,Naïve PAINE 在多个提示词数据集上优于现有方法。

原文摘要 · Abstract (English)

Text-to-Image (T2I) generation is primarily driven by Diffusion Models (DM) which rely on random Gaussian noise. Thus, like playing the slots at a casino, a DM will produce different results given the same user-defined inputs. This imposes a gambler's burden: To perform multiple generation cycles to obtain a satisfactory result. However, even though DMs use stochastic sampling to seed generation, the distribution of generated content quality highly depends on the prompt and the generative ability of a DM with respect to it. To account for this, we propose Naïve PAINE for improving the generative quality of Diffusion Models by leveraging T2I preference benchmarks. We directly predict the numerical quality of an image from the initial noise and given prompt. Naïve PAINE then selects a handful of quality noises and forwards them to the DM for generation. Further, Naïve PAINE provides feedback on the DM generative quality given the prompt and is lightweight enough to seamlessly fit into existing DM pipelines. Experimental results demonstrate that Naïve PAINE outperforms existing approaches on several prompt corpus benchmarks.

文生图扩散模型提示优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。