解决视频生成训练与推理不一致问题,提升长序列生成质量。
Stream Forcing: Constructing Unified Training Trajectory for Robust Streaming Video Generation

- 将采样过程建模为噪声水平的帧索引随机过程,统一训练轨迹。
- 在UCF-101上实现36.6%的FVD降低,零样本外推提升27.9%。
- 适合追求高效、稳定长视频生成的研究者和开发者。
流式视频生成在世界建模中具有巨大潜力,需在线连续推断未来帧以形成连续视频流。然而,流式视频扩散模型存在训练-推理不一致的根本问题:推理采用特定去噪顺序,而先进训练策略通常需要多样化的噪声水平配置。为解决训练覆盖与推理一致性之间的权衡,我们重新将视频扩散采样建模为噪声水平上的帧索引随机过程。在此随机过程空间中,构建一条从独立采样到推理一致采样的连续训练轨迹。进一步提出联合校准算法与时间相关采样算法,确保轨迹平滑性与跨帧相关性。基于此设计,提出Stream Forcing统一训练框架,平衡训练充分性与推理效率。大量实验表明,Stream Forcing在UCF-101基准上显著提升生成质量,FVD改善达36.6%;同时支持鲁棒的零样本长时序视频生成,FVD提升27.9%。
原文摘要 · Abstract (English)
Streaming video generation holds strong potential for world modeling, where future frames must be inferred online sequentially to form a continuous video stream. However, streaming video diffusion models introduce a fundamental train-inference mismatch: inference follows a specialized denoising order, whereas advanced training strategies typically require diverse noise-level configurations. To address this trade-off between train-inference consistency and training coverage, we reformulate the video diffusion sampling as a frame-indexed stochastic process over noise levels. Within this stochastic process space, we construct a continuous training trajectory along which the sampling schedule progressively evolves from independent sampling to inference-consistent sampling. We further introduce a joint calibration algorithm and a temporal correlative sampling algorithm to ensure trajectory smoothness and cross-frame correlation. Building on these designs, we propose Stream Forcing, a unified training framework for streaming video generation that balances training sufficiency and inference efficiency. Extensive experiments demonstrate that Stream Forcing significantly improves generation quality with a 36.6% FVD improvement on the UCF-101 benchmark. Furthermore, our method facilitates robust zero-shot extrapolation to long-horizon video generation with a 27.9% FVD improvement on the UCF-101 benchmark.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。