arXiv:2509.06942cs.AIcs.LG2025-09被引 29

通过在线调整奖励与插值恢复,实现扩散模型高效美学对齐

Directly Aligning the Full Diffusion Trajectory with Fine-Grained Human Preference

  • 用插值预定义噪声先验,避免多步去噪计算开销
  • 在线调整奖励使真实感与美感提升超3倍
  • 适合需要快速美学优化的生成系统开发者

近期研究证明,通过可微分奖励直接对齐扩散模型与人类偏好有效。但存在两大挑战:一是依赖多步去噪和梯度计算进行奖励评分,计算成本高,仅限少数扩散步骤优化;二是常需持续离线微调奖励模型以获得理想美学质量,如写实性或精确光照效果。为此,我们提出Direct-Align方法,通过预定义噪声先验,利用扩散状态是噪声与目标图像之间的插值这一特性,从任意时间步高效恢复原始图像,有效避免晚期时间步过度优化。同时引入语义相对偏好优化(SRPO),将奖励建模为文本条件信号,支持通过对正负提示词增强实现在线奖励调整,降低对离线奖励微调的依赖。在优化去噪过程与在线奖励调整的基础上,对FLUX模型进行微调,其人类评估的真实感与美学质量提升超过3倍。

原文摘要 · Abstract (English)

Recent studies have demonstrated the effectiveness of directly aligning diffusion models with human preferences using differentiable reward. However, they exhibit two primary challenges: (1) they rely on multistep denoising with gradient computation for reward scoring, which is computationally expensive, thus restricting optimization to only a few diffusion steps; (2) they often need continuous offline adaptation of reward models in order to achieve desired aesthetic quality, such as photorealism or precise lighting effects. To address the limitation of multistep denoising, we propose Direct-Align, a method that predefines a noise prior to effectively recover original images from any time steps via interpolation, leveraging the equation that diffusion states are interpolations between noise and target images, which effectively avoids over-optimization in late timesteps. Furthermore, we introduce Semantic Relative Preference Optimization (SRPO), in which rewards are formulated as text-conditioned signals. This approach enables online adjustment of rewards in response to positive and negative prompt augmentation, thereby reducing the reliance on offline reward fine-tuning. By fine-tuning the FLUX model with optimized denoising and online reward adjustment, we improve its human-evaluated realism and aesthetic quality by over 3x.

扩散模型偏好对齐在线优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。