用奖励信号指导生成,实现高效高质量自回归视频合成
Reward-Forcing: Autoregressive Video Generation with Reward Feedback
- 用奖励反馈替代教师模型引导自回归视频生成
- 在VBench上取得84.92分,超越部分双向模型
- 训练更简单,适合追求实时生成的场景
尽管多数先前的视频生成工作依赖于双向架构,近期研究尝试将这些模型转化为自回归变体以支持近实时生成。然而,这类转化往往高度依赖教师模型,尤其在缺乏强自回归教师时,性能受限,输出质量通常落后于双向模型。本文提出一种新方法,利用奖励信号引导生成过程,实现更高效、可扩展的自回归视频生成。该方法简化了训练流程,同时保持高视觉保真度和时间一致性。在标准基准上的大量实验表明,本方法表现与现有自回归模型相当,某些情况下甚至超过同等规模的双向模型,避免了教师架构带来的约束。例如,在VBench上,本方法总得分为84.92,接近顶尖自回归方法的84.31分,但无需复杂的异构蒸馏。
原文摘要 · Abstract (English)
While most prior work in video generation relies on bidirectional architectures, recent efforts have sought to adapt these models into autoregressive variants to support near real-time generation. However, such adaptations often depend heavily on teacher models, which can limit performance, particularly in the absence of a strong autoregressive teacher, resulting in output quality that typically lags behind their bidirectional counterparts. In this paper, we explore an alternative approach that uses reward signals to guide the generation process, enabling more efficient and scalable autoregressive generation. By using reward signals to guide the model, our method simplifies training while preserving high visual fidelity and temporal consistency. Through extensive experiments on standard benchmarks, we find that our approach performs comparably to existing autoregressive models and, in some cases, surpasses similarly sized bidirectional models by avoiding constraints imposed by teacher architectures. For example, on VBench, our method achieves a total score of 84.92, closely matching state-of-the-art autoregressive methods that score 84.31 but require significant heterogeneous distillation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。