揭示扩散模型奖励引导中的偏差根源并提出修正方案
Are we really tilting? The mechanics of reward guidance in flow and diffusion models

- 发现奖励引导偏差源于有限粒子估计的近似误差
- 理论证明存在模式内过优化和模式选择失败两种故障
- 提出无需额外计算的奖励衰减策略,适配图像生成场景
奖励引导算法在推理时将已训练的生成过程导向奖励倾斜分布。尽管实证效果显著,但此类方法易出现奖励滥用:模型为最大化奖励而牺牲对原始分布的保真度。以往研究归因于神经奖励函数的复杂性或扩散训练中的隐式偏见,但其根本成因仍不明确。本文表明,奖励滥用源于大多数实际实现中采用的奖励引导扩散方法——基于有限粒子的插值估计(Doob h-函数)所产生的近似误差,即使在最简单的高斯与高斯混合目标、二次奖励设置下亦然。我们以闭式解形式,分离出该估计器的两种独立失效模式:导致各模式内部过度优化,且无法选择高奖励模式。为此,我们提出一种闭式奖励衰减调度策略,可无额外计算修正模式内偏差;同时阐明了 best-of-n 采样在弥补模式选择失败中的作用。在高斯混合目标、二维棋盘图及 FLUX.1 文生图任务上的实验验证了理论结论在实际场景中的有效性。
原文摘要 · Abstract (English)
Reward guidance algorithms steer a learned generative process toward the reward-tilted measure at inference time. While empirically powerful, these methods are prone to reward hacking: the guided model over-optimizes the reward at the cost of fidelity to the learned distribution. Prior work has attributed this to the complexity of neural reward functions or implicit biases in diffusion training, but its fundamental origins remain poorly understood. We show that reward hacking arises from an approximation made in most practical implementations of reward-guided diffusion -- finite-particle plug-in estimation of the Doob h-function -- even in the simplest non-trivial settings of Gaussian and Gaussian mixture targets with quadratic rewards. In closed form, we isolate two distinct failure modes of the plug-in estimator: it leads to reward hacking within each mode and it cannot select high-reward modes. We propose a closed-form reward damping schedule that corrects the within-mode bias with no additional compute, and clarify the role of best-of-n sampling in compensating for the mode selection failure. Experiments on Gaussian mixture targets, a 2D checkerboard, and FLUX.1 text-to-image generation confirm that our theoretical insights carry over to practical settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。