提出新方法让一键生成图像更精准,避免失真。
Diff-Instruct with Diffused Reward: Towards Principled One-step Generator RL

- 通过积分KL最小化构建轨迹对齐框架,统一奖励与生成过程。
- 在单步生成中实现优于50步教师模型的偏好对齐效果。
- 适合追求高效高质图像生成的研究者与开发者。
最近的一键文本到图像生成技术实现了实时合成,兼具高效与高质量。此前的强化学习方法将图像空间奖励优化与扩散噪声空间分布匹配结合,但存在终端奖励优化与生成动态不匹配的问题,导致优化常利用随机自由度,在提升奖励的同时损害图像保真度。为此,我们提出基于积分KL最小化的无数据轨迹级对齐框架Diff-Instruct with Diffused Reward (DIDR),将RLHF最优的奖励倾斜清洁图像分布沿扩散轨迹传播至所有噪声水平。该目标与清洁图像RLHF具有相同最小值,同时自然诱导出扩散奖励得分(DRS),作为奖励驱动的参考得分函数修正项。为实用化,进一步引入可微短步去噪实现的扩散奖励代理(DRP),高效估计DRS。大量实验表明,DIDR在单步生成中持续优于现有SDXL基线。当迁移至6B DiT骨干网络(Z-Image)时,其在仅需单步生成的情况下,超越50步教师模型的偏好对齐性能。
原文摘要 · Abstract (English)
Recent advances in one-step text-to-image generation have enabled real-time synthesis with remarkable efficiency and quality. Previous reinforcement learning methods for one-step generators combine image-space reward optimization with diffusion noisy-space distribution matching. This paradigm brings challenges due to a mismatch between terminal reward optimization and the underlying generative dynamics. As a result, optimization tends to exploit stochastic degrees of freedom, often improving reward at the expense of image fidelity. To address this issue, we propose Diff-Instruct with Diffused Reward (DIDR), a data-free trajectory-level alignment framework derived from Integral KL minimization. DIDR propagates the RLHF-optimal reward-tilted clean-image distribution across all noise levels along the diffusion trajectory. We show that this objective admits the same minimizer as clean-image RLHF, while naturally inducing the Diffused Reward Score (DRS), which acts as a reward-driven correction to the reference score function. To make this practical, we further introduce the Diffused Reward Proxy (DRP), an efficient estimator of DRS based on differentiable short-step denoising. Extensive experiments demonstrate that DIDR consistently Pareto-dominates existing one-step SDXL baselines. Moreover, when transferred to a 6B DiT backbone (Z-Image), DIDR surpasses its 50-step teacher in preference alignment while requiring only a single generation step.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。