提出LIVE模型,实现长时视频生成误差可控。
LIVE: Long-horizon Interactive Video World Modeling
- 用循环一致性约束控制误差积累,无需教师模型。
- 在长序列生成中保持视觉质量,超越训练长度。
- 适合需要稳定长时视频生成的研究与应用。
自回归视频世界模型根据动作预测未来视觉观测,但在长时程生成中常因误差累积而失效。现有方法依赖预训练教师模型和序列级分布匹配,增加计算开销且无法阻止训练外的误差传播。本文提出LIVE,一种通过新颖循环一致性目标实现误差有限累积的长时交互式视频世界模型,无需教师蒸馏。LIVE先从真实帧进行前向滚动,再执行反向生成以重建初始状态,对重建终端状态计算扩散损失,从而显式约束长时程误差传播。此外,我们提供统一视角并引入渐进式训练课程以稳定训练。实验表明,LIVE在长时程基准上达到最先进性能,能生成远超训练长度、高质量且稳定的视频。
原文摘要 · Abstract (English)
Autoregressive video world models predict future visual observations conditioned on actions. While effective over short horizons, these models often struggle with long-horizon generation, as small prediction errors accumulate over time. Prior methods alleviate this by introducing pre-trained teacher models and sequence-level distribution matching, which incur additional computational cost and fail to prevent error propagation beyond the training horizon. In this work, we propose LIVE, a Long-horizon Interactive Video world modEl that enforces bounded error accumulation via a novel cycle-consistency objective, thereby eliminating the need for teacher-based distillation. Specifically, LIVE first performs a forward rollout from ground-truth frames and then applies a reverse generation process to reconstruct the initial state. The diffusion loss is subsequently computed on the reconstructed terminal state, providing an explicit constraint on long-horizon error propagation. Moreover, we provide an unified view that encompasses different approaches and introduce progressive training curriculum to stabilize training. Experiments demonstrate that LIVE achieves state-of-the-art performance on long-horizon benchmarks, generating stable, high-quality videos far beyond training rollout lengths.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。