提出RAVEN模型,提升视频实时生成的长期一致性与质量。
RAVEN: Real-time Autoregressive Video Extrapolation with Consistency-model GRPO

- 将自回归视频生成重构成交错的历史与去噪状态序列,对齐训练与推理分布。
- 在多个评估指标上超越现有因果视频蒸馏方法,长时生成更稳定。
- 结合一致性模型强化学习,无需额外扩散过程,优化生成策略。
因果自回归视频扩散模型通过从已生成内容中外推未来片段实现实时流式生成。从高保真双向教师模型中蒸馏此类生成器可获得性能优异的少步模型,但训练时遇到的历史分布与推理时实际分布之间的持续差异,限制了长时程生成质量。我们提出实时自回归视频外推网络(RAVEN),一种训练-测试框架,将每个自回滚过程重新组织为清洁历史端点与噪声去噪状态的交错序列。该形式使训练注意力与推理时的外推对齐,并允许下游块损失监督未来预测所依赖的历史表示。我们进一步提出一致性模型组相对策略优化(CM-GRPO),将一致性采样步骤重构为条件高斯转移,并直接在此核上应用在线强化学习(RL),避免了先前流模型强化学习中采用的欧拉-马鲁亚姆辅助过程。实验表明,RAVEN在质量、语义和动态度评估上均优于近期因果视频蒸馏基线,且与CM-GRPO结合后进一步提升性能。
原文摘要 · Abstract (English)
Causal autoregressive video diffusion models support real-time streaming generation by extrapolating future chunks from previously generated content. Distilling such generators from high-fidelity bidirectional teachers yields competitive few-step models, yet a persistent gap between the history distributions encountered during training and those arising at inference constrains generation quality over long horizons. We introduce the Real-time Autoregressive Video Extrapolation Network (RAVEN), a training-time test framework that repacks each self rollout into an interleaved sequence of clean historical endpoints and noisy denoising states. This formulation aligns training attention with inference-time extrapolation and allows downstream chunk losses to supervise the history representations on which future predictions depend. We further propose Consistency-model Group Relative Policy Optimization (CM-GRPO), which reformulates a consistency sampling step as a conditional Gaussian transition and applies online Reinforcement Learning (RL) directly to this kernel, avoiding the Euler-Maruyama auxiliary process adopted in prior flow-model RL formulations. Experiments demonstrate that RAVEN surpasses recent causal video distillation baselines across quality, semantic, and dynamic degree evaluations, and that CM-GRPO provides further gains when combined with RAVEN.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。