arXiv:2604.19234cs.CV2026-04被引 1

让扩散模型生成更准:按阶段分配奖励,提升图文一致性和运动连贯性

Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation

论文配图:Learning to Credit the Right Steps: Objective-aware Process Optimization for Visual Generation
图 1 · 摘自论文原文
  • 按不同去噪阶段动态分配奖励权重,避免统一信号干扰
  • 在图像和视频生成任务中,各项指标均显著提升
  • 适合追求高质量生成结果的视觉生成研究者

强化学习,尤其是组相对策略优化(GRPO),已成为利用人类偏好信号对后训练视觉生成模型进行优化的有效框架。然而,其效果受限于粗粒度的奖励信用分配。在现代视觉生成中,常使用多个奖励模型来捕捉异质目标,如视觉质量、运动一致性与文本对齐。现有GRPO流程通常将这些奖励合并为单一静态标量,并在整个扩散轨迹中均匀传播,忽略了不同去噪步骤的阶段性作用,导致优化信号错时或不兼容。为此,我们提出目标感知轨迹信用分配(OTCA),一种细粒度的GRPO训练结构化框架。OTCA包含两个核心组件:轨迹级信用分解用于估计各去噪步骤的相对重要性;多目标信用分配则在去噪过程中自适应加权并融合多个奖励信号。通过联合建模时间信用与目标级信用,OTCA将粗粒度奖励监督转化为结构化、时步感知的训练信号,更契合扩散生成的迭代特性。大量实验表明,OTCA在图像与视频生成质量上均实现持续提升,覆盖多种评估指标。

原文摘要 · Abstract (English)

Reinforcement learning, particularly Group Relative Policy Optimization (GRPO), has emerged as an effective framework for post-training visual generative models with human preference signals. However, its effectiveness is fundamentally limited by coarse reward credit assignment. In modern visual generation, multiple reward models are often used to capture heterogeneous objectives, such as visual quality, motion consistency, and text alignment. Existing GRPO pipelines typically collapse these rewards into a single static scalar and propagate it uniformly across the entire diffusion trajectory. This design ignores the stage-specific roles of different denoising steps and produces mistimed or incompatible optimization signals. To address this issue, we propose Objective-aware Trajectory Credit Assignment (OTCA), a structured framework for fine-grained GRPO training. OTCA consists of two key components. Trajectory-Level Credit Decomposition estimates the relative importance of different denoising steps. Multi-Objective Credit Allocation adaptively weights and combines multiple reward signals throughout the denoising process. By jointly modeling temporal credit and objective-level credit, OTCA converts coarse reward supervision into a structured, timestep-aware training signal that better matches the iterative nature of diffusion-based generation. Extensive experiments show that OTCA consistently improves both image and video generation quality across evaluation metrics.

扩散模型强化学习生成质量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。