arXiv:2607.08020cs.CV2026-07

提出SAGA方法,解决自回归视频生成中的抖动与漂移问题。

SAGA: Stable Acceleration Guidance for Autoregressive Video Generation

论文配图:SAGA: Stable Acceleration Guidance for Autoregressive Video Generation
图 1 · 摘自论文原文
  • 基于离散潜变量加速度检测高频时间扰动,设计无训练引导机制。
  • 在Self-Forcing模型上提升时序质量至97.91,图像质量达70.51。
  • 无需重训练,适配现有自回归扩散模型,适合高质量视频生成场景。

自回归视频扩散模型虽支持高效流式传输和长时序生成,但反复使用生成的潜变量作为因果上下文会放大时间误差,导致闪烁、运动抖动和结构漂移。本文从谱运动学角度分析此问题,发现离散潜变量加速度是揭示不稳定性高频时间扰动的有效信号。为此,提出SAGA——一种无需训练的稳定加速度引导方法。SAGA结合基于有限窗Slepian投影的加速度域谱引导目标,以及抑制短程时间相关性但保留长程运动结构的自回归噪声初始化策略。无需修改主干模型,SAGA可直接应用于主流分块自回归扩散模型。大量实验表明,SAGA在多个模型上持续提升时序质量。在Self-Forcing模型上,时序质量从97.30提升至97.91,图像质量从69.60提升至70.51。谱分析与人工偏好研究进一步证实,SAGA有效降低时间不稳定性,同时保持视觉保真度。

原文摘要 · Abstract (English)

Autoregressive video diffusion enables efficient streaming and long-horizon video generation, but repeatedly reusing generated latents as causal context can amplify temporal errors, resulting in flickering, motion jitter, and structural drift. In this paper, we investigate this failure mode from a spectral kinematic perspective and identify discrete latent acceleration as an effective signal for revealing unstable high-frequency temporal perturbations. To this end, we propose SAGA, a training-free \textbf{\textit{s}}table \textbf{\textit{a}}cceleration \textbf{\textit{g}}uidance approach for \textbf{\textit{a}}utoregressive video generation. SAGA integrates an acceleration domain spectral guidance objective based on finite-window Slepian projections with a structured autoregressive noise initialization strategy that suppresses short-range temporal correlations while preserving long-range motion structure. Without retraining or modifying the backbone, SAGA can be directly applied to existing chunk-wise autoregressive diffusion models, which is the prevalent setting for high-quality generation. Extensive experiments show that SAGA consistently improves temporal quality across multiple autoregressive diffusion models. On Self-Forcing, SAGA improves Temporal Quality from 97.30 to 97.91 and Image Quality from 69.60 to 70.51. Moreover, spectral analysis and human preference studies demonstrate that SAGA reduces temporal instability while maintaining visual fidelity.

视频生成扩散模型自回归时序稳定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。