提出One-Forcing方法,实现高效高质量的单步自回归视频生成。
One-Forcing: Towards Stable One-Step Autoregressive Video Generation

- 在DMD目标上加入辅助GAN损失,提升生成稳定性与质量。
- 单步生成在VBench上得分83.76,超越现有单步方法并接近多步模型。
- 训练成本仅为分块模型的1/3,可稳定实现单帧自回归生成。
近期进展显著提升了自回归框架下的实时交互式视频生成能力。然而,大多数现有少步自回归视频生成方法通常从多步教师模型蒸馏而来,默认采用4步采样配置,部署时仍存在显著延迟,且进一步减少采样步数时质量严重下降,尤其在单步设置下表现不佳。轨迹一致性蒸馏方法常导致动态性弱,而基于DMD的方法(如Self-Forcing)则易生成模糊帧。为此,我们提出One-Forcing,一种简单但有效的方案,通过在DMD目标中引入辅助GAN损失,实现高质量、高效率的单步视频生成。在VBench上的实验表明,One-Forcing取得83.76的总分,成为当前单步因果视频生成方法中的最佳表现,且仍可与强大多步方法媲美。此外,我们进一步证明,仅需分块模型三分之一的训练成本,即可稳定实现单步帧级自回归生成,这一设置此前方法未能成功达成。
原文摘要 · Abstract (English)
Recent advances have substantially improved real-time interactive video generation in the autoregressive regime. However, most existing few-step autoregressive video generation methods, often distilled from a corresponding many-step teacher, default to a 4-step sampling configuration, which still incurs considerable latency during deployment and suffers from severe quality degradation when the number of sampling steps is further reduced, particularly in the one-step setting. Trajectory-style consistency distillation methods often produce videos with weak dynamics, while DMD-based approaches, such as Self-Forcing, tend to yield blurry frames. To address this challenge, we propose One-Forcing, a simple yet effective approach which augments the DMD objective with an auxiliary GAN loss for high-quality and efficient one-step video generation. Experiments on VBench show that One-Forcing achieves a total score of 83.76, establishing state-of-the-art performance among one-step causal video generation methods and remaining competitive with strong many-step approaches. We further demonstrate that one-step framewise autoregressive generation can be achieved stably with merely one-third of the training cost of the chunkwise model, a setting that prior methods have failed to achieve successfully.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。