arXiv:2511.18742cs.CVcs.AI2025-11被引 1

用反向离散化提升文本生成图像效率,更省算力且更符合人类偏好。

ProxT2I: Efficient Reward-Guided Text-to-Image Generation via Proximal Diffusion

  • 改用反向离散和可学习的近似算子替代传统得分函数
  • 在1500万高质量人脸数据上训练,采样效率显著提升
  • 适合追求轻量化与高拟真图像生成的开发者

扩散模型已成为多种生成任务的主流范式,包括提示条件生成。现有采样器多依赖逆扩散过程的前向离散化及从数据中学习的得分函数,此类方法往往耗时且不稳定,需大量采样步骤才能生成高质量图像。本文提出基于反向离散化的文本到图像(T2I)扩散模型ProxT2I,采用学习得到的条件近似算子代替得分函数,并结合强化学习与策略优化技术,使采样器能针对特定任务奖励进行优化。此外,我们构建了一个包含1500万张高质量真人图像及细粒度描述的开源数据集LAION-Face-T2I-15M,用于训练与评估。相比基于得分函数的基线方法,本方法在采样效率和人类偏好对齐方面均有提升,在保持与现有先进开源T2I模型相当性能的同时,所需计算资源更少、模型更小,为人类图像生成提供了一种轻量高效的新方案。

原文摘要 · Abstract (English)

Diffusion models have emerged as a dominant paradigm for generative modeling across a wide range of domains, including prompt-conditional generation. The vast majority of samplers, however, rely on forward discretization of the reverse diffusion process and use score functions that are learned from data. Such forward and explicit discretizations can be slow and unstable, requiring a large number of sampling steps to produce good-quality samples. In this work we develop a text-to-image (T2I) diffusion model based on backward discretizations, dubbed ProxT2I, relying on learned and conditional proximal operators instead of score functions. We further leverage recent advances in reinforcement learning and policy optimization to optimize our samplers for task-specific rewards. Additionally, we develop a new large-scale and open-source dataset comprising 15 million high-quality human images with fine-grained captions, called LAION-Face-T2I-15M, for training and evaluation. Our approach consistently enhances sampling efficiency and human-preference alignment compared to score-based baselines, and achieves results on par with existing state-of-the-art and open-source text-to-image models while requiring lower compute and smaller model size, offering a lightweight yet performant solution for human text-to-image generation.

文本生成图像扩散模型采样效率轻量化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。