用分布匹配统一视频生成对齐与蒸馏,提升质量与偏好契合度。
Joint Alignment and Distillation for Video Generation via Sample-Guided Distribution Matching

- 基于分布匹配框架,同步优化生成质量和人类偏好
- 在多个视频模型上实现比独立或分步方法更好的性能
- 无需强化学习的多步评估,计算效率更高
视频生成模型对齐人类偏好主要依赖强化学习(RL),但存在计算开销大等问题。现有流程通常将强化学习与蒸馏视为分离阶段:先进行强化学习会导致高昂计算成本,而先蒸馏再强化常引发模型坍塌。为此,我们提出一种基于分布匹配(DM)的统一单阶段优化框架。标准分布匹配通过最小化真实与虚假模型间的差距来更新模型,提升生成清晰度与保真度。在此基础上,我们引入DM-Align,利用偏好对或组内探索所形成的分布差距,直接构建引导模型向人类偏好的样本靠拢的互补梯度方向。通过协同两个梯度方向,该方法避免了传统强化学习中多步奖励评估和复杂的ODE-SDE转换。在多个基础视频生成模型上的实验表明,该样本引导框架能稳健提升蒸馏质量与偏好对齐效果,持续优于单一变体及顺序两阶段流程。
原文摘要 · Abstract (English)
Aligning video generative models to human preferences heavily relies on Reinforcement Learning (RL), which suffers from extensive computational overhead. Existing workflows typically treat RL and distillation as disconnected stages: applying RL before distillation incurs prohibitive computational costs, whereas applying RL after distillation frequently leads to model collapse. To overcome these limitations, we propose a unified, single-stage optimization framework grounded in Distribution Matching (DM). In the standard DM framework, distillation updates the model via a gradient direction that minimizes the gap between the real and fake models, guiding generations toward clarity and high fidelity. Building upon this, we introduce DM-Align, which derives a complementary gradient direction to guide the model toward human-preferred samples. Inspired by DPO and GRPO, our method leverages the distributional gap -- formulated from either preference pairs or intra-group exploration -- to directly construct this preference-guided gradient. By synergizing these two gradient directions, our approach eliminates the need for multi-step reward evaluation and complex ODE-SDE conversions inherent in traditional RL. Comprehensive experiments across multiple foundational video models demonstrate that this sample-guided framework robustly enhances both distillation quality and preference alignment, consistently outperforming both standalone variants and sequential two-stage pipelines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。