通过渐进式噪声上下文提升视频生成的时序一致性与速度
In-Context Forcing: Uncovering Context Effects in Autoregressive Video Diffusion

- 采用噪声递减的上下文策略,动态调节帧间依赖强度
- 在VBench上实现视觉质量与推理速度双突破
- 支持跨帧并行去噪,加速推理且不损失性能
当前少步自回归视频扩散模型依赖前一帧完全去噪的清晰帧作为上下文,但这些清晰帧泄露过多局部细节,导致模型走捷径,削弱时序语义与动态表现。受扩散过程即掩码视角启发,我们研究了噪声上下文对生成的影响。然而,统一噪声水平的上下文提供引导不足,造成时序不一致。为此,提出In-Context Forcing:一种渐进式自回归范式,利用噪声水平逐级下降的上下文。通过减少远距离帧的掩码、增加邻近帧的掩码,实现自适应引导,有效保障时序一致性与高帧间动态性。此外,解耦对前序清晰帧的强依赖,支持跨帧并行去噪,在不牺牲性能的前提下显著加速推理。在VBench上的大量实验表明,本方法在视觉保真度和推理速度上均显著优于现有最佳方案。
原文摘要 · Abstract (English)
Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame. However, these clean frames leak excessive local details, which causes the model to take shortcuts, resulting in compromised temporal semantics and dynamics. Inspired by the perspective of diffusion as masking, we explore the impact of noisy contexts on few-step autoregressive generation. Yet, simply applying contexts with the same noise levels provides insufficient guidance, leading to poor temporal consistency. To resolve this dilemma, we introduce In-Context Forcing, a progressive autoregressive paradigm that utilizes contexts with decreasing noise levels. By applying less masking to distant frames and more masking to adjacent ones, this approach provides adaptive guidance, effectively ensuring both robust temporal consistency and high inter-frame dynamics. Furthermore, by decoupling the strict dependence on previous clean frames, our paradigm enables cross-frame parallel denoising, achieving substantial inference acceleration without sacrificing performance. Extensive experiments on VBench demonstrate that our method significantly outperforms state-of-the-art approaches in both visual fidelity and inference speed.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。