arXiv:2607.27036cs.CVcs.LG2026-07

发现视频生成误差累积源于隐空间维度坍缩,提出轻量正则化稳定生成。

Mitigating Compounding Error via Video Representation Regularization

论文配图:Mitigating Compounding Error via Video Representation Regularization
图 1 · 摘自论文原文
  • 通过分析模型内部表示,发现误差累积与隐空间有效秩下降相关。
  • 在VBench上,美学质量提升至55.56,图像质量达72.08,显著优于基线。
  • 适合关注长时序视频生成稳定性的研究者和工业应用开发者。

基于扩散的视频世界模型支持机器人、自动驾驶与仿真任务中的长序列自回归视频生成,但滑动窗口推理存在严重误差累积,导致帧质量随时间退化。尽管此现象广泛存在,其内在机制及长期生成稳定性问题仍未解决。本文研究视频世界模型的内部表示动态,发现误差累积与隐藏表示的维度坍缩密切相关:生成漂移初期,模型表示的有效秩急剧下降,表明表征退化与长期滚动不稳定性紧密关联。此外,我们发现单纯的数据规模扩展无法提升模型抗误差漂移能力,挑战主流扩展范式。为此,提出视频表示正则化方法,一种轻量级训练约束,可稳定潜在表示并抑制迭代误差积累。相比Diffusion Forcing,本方法在VBench的美学质量指标上从38.65提升至55.56,在图像质量上从44.37提升至72.08。本工作首次建立自回归视频漂移与模型内部表示之间的联系,采用erank作为误差累积的量化指标,揭示视频世界模型的反直觉扩展局限,并提出一种简单有效的正则化策略以增强长视频生成鲁棒性。

原文摘要 · Abstract (English)

Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time. Although this phenomenon has been widely observed, the underlying mechanism of compounding error and how to achieve stable long-horizon generation remain largely unresolved. In this paper, we investigate the internal representation dynamics of video world models and discover that compounding error is tightly coupled with dimensional collapse of hidden representations. Specifically, the effective rank of model representations sharply decreases at the onset of generation drift, revealing a strong connection between representational degradation and long-term rollout instability. Furthermore, we find that pure training data scaling fails to boost model resistance to error drift, contradicting mainstream scaling paradigms. To address this problem, we propose video representation regularization, a lightweight training constraint that stabilizes latent representations and suppresses iterative error accumulation. Compared with Diffusion Forcing, our method achieves improvements from 38.65 to 55.56 and from 44.37 to 72.08 on the Aesthetic Quality and Imaging Quality metrics of VBench. Our work establishes the first connection between autoregressive video drifting and model internal representations, adopts erank as a quantitative metric for error accumulation, reveals counterintuitive scaling limitations for video world models, and presents a simple yet effective regularization strategy to improve long video generation robustness.

视频生成扩散模型误差累积表示正则化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。