arXiv:2503.10704cs.CVcs.MM2025-03被引 20

提出统一框架分析自回归视频生成模型的遗忘与退化问题

Error Analyses of Auto-Regressive Video Diffusion Models: A Unified Framework

  • 基于条件互信息建模历史遗忘,证明多帧输入可缓解该问题
  • 发现帧质量随时间衰减与累积误差相关,可预测不同调度效果
  • 设计新评测任务揭示两类错误存在强关联,适用于长视频生成研究

自回归视频扩散模型(AR-VDMs)在生成长时、逼真视频方面表现优异,但面临两大缺陷:历史遗忘(模型逐渐丢失已生成内容)和时间退化(帧质量随生成时间下降)。现有研究缺乏理论分析,实证理解也较薄弱。本文提出Meta-ARVDM统一分析框架,通过共享的自回归结构同时研究两类错误。我们证明历史遗忘由生成输出与前序帧的条件互信息刻画,且增加过去帧数可单调缓解该问题,从理论上支持了现有实践中的直觉。此外,理论表明标准评估指标无法捕捉此现象,因此提出基于“针在草堆中”任务的新评测协议,应用于DMLab和Minecraft封闭环境。进一步证明时间退化可由每步误差的累积和量化,无需实际生成视频即可预测不同调度器的表现。最终实验发现历史遗忘与时间退化存在强经验相关性,这一关联此前未被报道。

原文摘要 · Abstract (English)

Auto-Regressive Video Diffusion Models (AR-VDMs) have shown strong capabilities in generating long, photorealistic videos, but suffer from two key limitations: (i) history forgetting, where the model loses track of previously generated content, and (ii) temporal degradation, where frame quality deteriorates over time. Yet a rigorous theoretical analysis of these phenomena is lacking, and existing empirical understanding remains insufficiently grounded. In this paper, we introduce Meta-ARVDM, a unified analytical framework that studies both errors through the shared autoregressive structure of AR-VDMs. We show that history forgetting is characterized by the conditional mutual information between the generated output and preceding frames, conditioned on inputs, and prove that incorporating more past frames monotonically alleviates history forgetting, thereby theoretically justifying a common belief in existing works. Moreover, our theory reveals that standard metrics fail to capture this effect, motivating a new evaluation protocol based on a ``needle-in-a-haystack'' task in closed-ended environments (DMLab and Minecraft). We further show that temporal degradation can be quantified by the cumulative sum of per-step errors, enabling prediction of degradation for different schedulers without video rollout. Finally, our evaluation uncovers a strong empirical correlation between history forgetting and temporal degradation, a connection not previously reported.

视频生成扩散模型错误分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。