arXiv:2606.05478cs.CVcs.LG2026-06

预测文生图生成前的人类偏好,可低成本提升图像质量。

Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So?

论文配图:Can We Predict The Human Preference For Text-to-Image Content Prior To Generation And Is It Even Useful To Do So?
图 1 · 摘自论文原文
  • 用初始噪声预测生成前的人类偏好得分。
  • 预测准确率高,且硬件开销几乎忽略不计。
  • 适合在小模型本地部署时优化生成效果。

扩散模型(DM)通过从用户提示生成高质量、逼真的视觉内容,革新了文本驱动的图像生成。与以往基于感知相似性指标(如FID、PSNR)的评估不同,DM推动了人类偏好度量(HPM)的发展,将人类判断量化为标量值。然而,DM的生成过程本质上是随机的,初始噪声直接影响输出质量,尤其在小型模型和本地部署场景中更为明显。本文首先探究在投入计算资源前,能否预测标量形式的HPM得分;进一步研究能否利用该预测提升生成图像质量,并分析哪些HPM更适合此任务。结果表明,不仅可行,且可实现几乎零硬件开销。

原文摘要 · Abstract (English)

Diffusion Models (DM) have revolutionized text-driven generation by enabling the synthesis of high-quality, photorealistic visual content from user prompts. Whereas prior advances in visual generation such as VAEs and GANs were primarily evaluated on perceptual or visual similarity metrics such as FID PSNR, DM advances have fostered the development of more advanced Human Preference Metrics (HPM) that model and quantify human judgment as scalar values. However, DMs synthesize content using an inherently stochastic process where random noise seeds generation. The initial random noise directly affects the quality of generated outputs, both qualitatively and quantitatively. This influence is pronounced in smaller models for local deployment scenarios. Given this phenomenon, we first investigate to what extent we can predict scalar HPM scores prior to committing compute resources for generation. Further, we then investigate to what extent we can leverage such prediction to improve the quality of generated images, and also study which HPMs are best suited for this task. Our investigation reveals that not only is this possible, but that it is feasible to achieve negligible hardware overhead.

文生图扩散模型人类偏好预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。