用奖励函数直接指导生成,让图像快速变真实。
Reward-Instruct: A Reward-Centric Approach to Fast Photo-Realistic Image Generation
- 不依赖复杂蒸馏损失,通过奖励调整参数分布来训练
- 在文本到图像生成中达到顶尖视觉质量和量化指标
- 适合追求高效、稳定且可控图像生成的研究者
本文解决高质快速图像生成与复杂人类偏好对齐的挑战。尽管扩散模型与蒸馏技术推动了生成加速,但如何有效融合奖励反馈以提升可控性和偏好对齐仍是难题。现有方法常依赖昂贵的扩散蒸馏损失,而本文发现:当条件更具体时,精心设计的奖励函数成为训练强少步生成模型的主要驱动力。为此提出Reward-Instruct,一种简单高效的奖励中心型方法,将预训练扩散模型转化为少步生成器。该方法无需复杂蒸馏损失,通过从奖励倾斜的参数分布中迭代采样更新模型,显著降低计算成本。实验表明,在文本到图像生成任务中,Reward-Instruct在视觉质量与定量指标上优于依赖蒸馏的方法,且对奖励函数选择更具鲁棒性。
原文摘要 · Abstract (English)
This paper addresses the challenge of achieving high-quality and fast image generation that aligns with complex human preferences. While recent advancements in diffusion models and distillation have enabled rapid generation, the effective integration of reward feedback for improved abilities like controllability and preference alignment remains a key open problem. Existing reward-guided post-training approaches targeting accelerated few-step generation often deem diffusion distillation losses indispensable. However, in this paper, we identify an interesting yet fundamental paradigm shift: as conditions become more specific, well-designed reward functions emerge as the primary driving force in training strong, few-step image generative models. Motivated by this insight, we introduce Reward-Instruct, a novel and surprisingly simple reward-centric approach for converting pre-trained base diffusion models into reward-enhanced few-step generators. Unlike existing methods, Reward-Instruct does not rely on expensive yet tricky diffusion distillation losses. Instead, it iteratively updates the few-step generator's parameters by directly sampling from a reward-tilted parameter distribution. Such a training approach entirely bypasses the need for expensive diffusion distillation losses, making it favorable to scale in high image resolutions. Despite its simplicity, Reward-Instruct yields surprisingly strong performance. Our extensive experiments on text-to-image generation have demonstrated that Reward-Instruct achieves state-of-the-art results in visual quality and quantitative metrics compared to distillation-reliant methods, while also exhibiting greater robustness to the choice of reward function.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。