arXiv:2410.18881cs.CVcs.AI2024-10中稿 · Transactions of Ma…被引 33

让一键生成图像模型更懂人类审美,效果领先。

Diff-Instruct++: Training One-step Text-to-image Generator Model to Align with Human Preferences

  • 用强化学习对齐人类偏好,不依赖图像数据。
  • 生成图像美学分达6.19,人类评分超28分。
  • 适合追求高效高质量图像生成的研究者。

一键式文本到图像生成模型具备推理速度快、架构灵活和生成性能优异等优势。本文首次研究此类模型与人类偏好的对齐问题。受基于人类反馈的强化学习(RLHF)成功启发,我们提出最大化期望人类奖励函数,并加入积分KL散度项以防止生成器发散。通过克服技术挑战,提出Diff-Instruct++(DI++),这是首个快速收敛且无需图像数据的人类偏好对齐方法。我们还揭示了使用条件控制生成(CFG)进行扩散蒸馏本质上是在执行带有DI++的RLHF,这一发现为未来研究提供了理论洞见。实验中,我们使用DI++对基于UNet和DiT的一键生成模型进行对齐,参考过程分别为Stable Diffusion 1.5和PixelArt-α。基于DiT的模型在COCO验证提示数据集上取得6.19的美学分和1.24的图像奖励,人类偏好评分(HPSv2.0)达28.48,超越Stable Diffusion XL、DMD2、SD-Turbo及PixelArt-α等开源模型。理论与实证均表明DI++是高效的人类偏好对齐方案。

原文摘要 · Abstract (English)

One-step text-to-image generator models offer advantages such as swift inference efficiency, flexible architectures, and state-of-the-art generation performance. In this paper, we study the problem of aligning one-step generator models with human preferences for the first time. Inspired by the success of reinforcement learning using human feedback (RLHF), we formulate the alignment problem as maximizing expected human reward functions while adding an Integral Kullback-Leibler divergence term to prevent the generator from diverging. By overcoming technical challenges, we introduce Diff-Instruct++ (DI++), the first, fast-converging and image data-free human preference alignment method for one-step text-to-image generators. We also introduce novel theoretical insights, showing that using CFG for diffusion distillation is secretly doing RLHF with DI++. Such an interesting finding brings understanding and potential contributions to future research involving CFG. In the experiment sections, we align both UNet-based and DiT-based one-step generators using DI++, which use the Stable Diffusion 1.5 and the PixelArt-$α$ as the reference diffusion processes. The resulting DiT-based one-step text-to-image model achieves a strong Aesthetic Score of 6.19 and an Image Reward of 1.24 on the COCO validation prompt dataset. It also achieves a leading Human preference Score (HPSv2.0) of 28.48, outperforming other open-sourced models such as Stable Diffusion XL, DMD2, SD-Turbo, as well as PixelArt-$α$. Both theoretical contributions and empirical evidence indicate that DI++ is a strong human-preference alignment approach for one-step text-to-image models. The homepage of the paper is https://github.com/pkulwj1994/diff_instruct_pp.

图像生成偏好对齐扩散模型强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。