让少步扩散模型用上人类评分等不可导奖励,显著提升生成质量。
TDM-R1: Reinforcing Few-Step Diffusion Models with Non-Differentiable Reward
- 分离奖励学习与生成器训练,支持不可导奖励信号。
- 仅用4步推断就超越100步基准,提升图像质量和偏好对齐。
- 适用于文本到图像生成,尤其适合追求高效生成的场景。
尽管少步生成模型已显著降低图像和视频生成成本,但针对少步模型的通用强化学习方法仍是一个未解难题。现有强化学习方法依赖可微分奖励模型的反向传播,无法使用人类二元偏好、物体数量等重要不可导奖励信号。为此,我们提出TDM-R1,一种基于领先少步模型轨迹分布匹配(TDM)的新强化学习范式。TDM-R1将学习过程解耦为代理奖励学习与生成器学习,并开发了沿确定性生成轨迹获取每步奖励信号的实用方法,形成统一的强化学习后训练方案,显著提升少步模型在通用奖励下的能力。我们在文本渲染、视觉质量与偏好对齐等多个任务上进行了广泛实验,结果表明TDM-R1在域内与域外指标上均达到当前最优强化学习性能。此外,TDM-R1还可有效扩展至最新强大的Z-Image模型,在仅4个非微分步(NFE)下持续优于其100-NFE及少步变体。
原文摘要 · Abstract (English)
While few-step generative models have enabled powerful image and video generation at significantly lower cost, generic reinforcement learning (RL) paradigms for few-step models remain an unsolved problem. Existing RL approaches for few-step diffusion models strongly rely on back-propagating through differentiable reward models, thereby excluding the majority of important real-world reward signals, e.g., non-differentiable rewards such as humans' binary likeness, object counts, etc. To properly incorporate non-differentiable rewards to improve few-step generative models, we introduce TDM-R1, a novel reinforcement learning paradigm built upon a leading few-step model, Trajectory Distribution Matching (TDM). TDM-R1 decouples the learning process into surrogate reward learning and generator learning. Furthermore, we developed practical methods to obtain per-step reward signals along the deterministic generation trajectory of TDM, resulting in a unified RL post-training method that significantly improves few-step models' ability with generic rewards. We conduct extensive experiments ranging from text-rendering, visual quality, and preference alignment. All results demonstrate that TDM-R1 is a powerful reinforcement learning paradigm for few-step text-to-image models, achieving state-of-the-art reinforcement learning performances on both in-domain and out-of-domain metrics. Furthermore, TDM-R1 also scales effectively to the recent strong Z-Image model, consistently outperforming both its 100-NFE and few-step variants with only 4 NFEs. Project page: https://github.com/Luo-Yihong/TDM-R1
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。