arXiv:2605.07503cs.CV2026-05

让视频生成模型更懂人话,通过同步训练与推理路径提升对齐效果。

Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers

论文配图:Diffusion-APO: Trajectory-Aware Direct Preference Alignment for Video Diffusion Transformers
图 1 · 摘自论文原文
  • 设计新算法同步训练噪声与推理去噪路径,增强梯度信号。
  • 在多个数据集上超越基线,保持生成质量的同时加速推理。
  • 无需依赖复杂奖励模型,适合资源受限的多阶段对齐任务。

高效对齐大规模视频扩散模型与人类意图,需要一种可扩展且轨迹感知的路径,以弥合训练噪声分布与实际推理轨迹之间的差异。现有方法如直接偏好优化(DPO)和组相对策略优化(GRPO)常受制于有偏的复杂奖励模型或次优的时间步采样。本文提出扩散-对齐偏好优化(Diffusion-APO),通过同步训练噪声与推理时去噪路径,最大化梯度信号有效性,解决这一错位问题。为实现该算法的实际应用,我们构建了一个统一、模块化的强化学习人类反馈框架,集成在线排序、半在线锚定、离线精修及蒸馏感知漂移校正。该框架支持在不同数据和计算约束下灵活进行多阶段偏好对齐,无需依赖基于标量奖励的策略梯度。大量实验表明,Diffusion-APO在视觉质量和指令遵循能力上持续优于标准基线,同时在模型加速过程中有效保持生成保真度,提供了一条稳健的端到端可扩展视频扩散对齐路径。

原文摘要 · Abstract (English)

Efficiently aligning large-scale video diffusion models with human intent requires a scalable and trajectory-aware pathway that bridges the inherent discrepancy between training noise distributions and practical inference trajectories. While existing paradigms such as Direct Preference Optimization (DPO) and Group Relative Policy Optimization (GRPO) attempt to address this, they are often hindered by either reliance on bias-prone, complex reward models or suboptimal timestep sampling. In this paper, we propose Diffusion-APO (Aligned Preference Optimization), a trajectory-aware algorithm that resolves this misalignment by synchronizing training noise with inference-time denoising paths to maximize gradient signal efficacy. To translate this algorithmic innovation into a practical solution, we introduce a unified and modular RLHF framework that integrates online ranking, half-online anchoring, offline refinement, and distillation-aware drift correction. This framework enables flexible, multi-stage preference alignment across diverse data and computational constraints without relying on scalar-reward-based policy gradients. Through extensive experiments, we demonstrate that Diffusion-APO consistently outperforms standard baselines in visual quality and instruction following, while effectively preserving generative fidelity during model acceleration, providing a robust, end-to-end pathway for scalable video diffusion alignment.

视频生成扩散模型强化学习对齐优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。