通过反思优化扩散模型生成,提升图像视频的语义准确性和真实感。
Step Back to Move Forward: Reflection-Aware Preference Optimization for Visual Generation

- 用反向扩散过程修正生成轨迹,引导潜空间走向更真实的分布。
- 在不增加推理开销的前提下,内化搜索探索的优势,显著减少奖励欺骗。
- 适用于文本到图像/视频生成,对模型架构无要求,易集成到现有流程。
扩散模型已成为现代视觉生成的主流范式,在文生图和文生视频任务中取得了显著进展。为进一步对齐生成模型与人类偏好,强化学习(RL)作为后训练策略展现出强大潜力。然而,现有的基于策略梯度的方法往往探索效率低下,易陷入局部最优,导致语义忠实度和视觉真实性下降。为此,我们提出反射感知的GRPO(RA-GRPO),一种面向扩散生成模型的新型强化学习偏好对齐框架。核心思想是通过引入‘回溯’反思来改进‘前进’生成。我们首先提出扩散反思(Diffusion Reflection),利用弱估计器反演扩散过程,修正中间采样轨迹,引导潜在状态向真实数据流形的高概率区域逼近。此外,我们设计了反事实路径合成(Counterfactual Path Synthesis),将这些修正后的轨迹隐式提炼进策略中,使模型在不增加推理开销的情况下内化基于搜索的探索优势。在文本到图像和文本到视频模型上的大量实验表明,RA-GRPO显著优于现有方法,尤其在缓解奖励欺骗和提升泛化能力方面表现突出。该方法保持架构无关性,可无缝集成至标准流水线,为稳定偏好对齐提供了有前景的方向。
原文摘要 · Abstract (English)
Diffusion models have become the mainstream paradigm for modern visual generation and have substantially advanced multimedia content synthesis, especially in text-to-image and text-to-video tasks. To further align such generative models with human preferences, reinforcement learning (RL) has recently shown strong potential as a post-training strategy. Nevertheless, existing policy gradient-based methods often explore inefficiently, making them vulnerable to local optima that may degrade semantic faithfulness and visual realism. To address these challenges, we present Reflection-Aware GRPO (RA-GRPO), a new RL-based preference alignment framework for diffusion generative models. The core idea is to improve "forward" generation by incorporating "backward" reflection during optimization. We first introduce Diffusion Reflection, which rectifies intermediate sampling trajectories by inverting the diffusion process with a weak estimator, guiding latent states toward higher-probability regions of the true data manifold. Furthermore, we introduce Counterfactual Path Synthesis to implicitly distill these rectified trajectories into the policy, enabling the model to internalize the benefits of search-based exploration without incurring inference-time overhead. Extensive experiments on T2I and T2V models demonstrate that RA-GRPO significantly outperforms existing methods, particularly in mitigating reward hacking and improving generalization. The method remains architecture-agnostic and integrates seamlessly with standard pipelines, suggesting a promising direction for stable preference alignment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。