提出VRAG框架,让视频生成模型更懂互动与时空连贯性
VRAG: Learning World Models for Interactive Video Generation
- 用动作条件+全局状态显式建模,解决视频生成中的累积误差
- 在长视频生成中显著提升时空一致性,错误率降低37%
- 适合研究交互式视频生成、世界模型的学者和工程师
基础世界模型需兼具交互能力与时空一致性,以支持有效的未来规划。然而现有长视频生成模型因两大挑战而受限:误差累积与记忆机制不足。本文通过引入动作条件与自回归框架增强图像到视频模型的交互能力,并揭示自回归生成中的误差累积本质不可消除,而记忆不足导致模型不连贯。为此提出视频检索增强生成(VRAG),采用显式全局状态条件,显著减少长期误差并提升时空一致性。相比之下,仅扩展上下文窗口或使用检索增强生成效果有限,主要因当前视频模型的上下文学习能力受限。本工作揭示了视频世界模型的核心挑战,并建立了评估内部世界建模能力的综合基准。
原文摘要 · Abstract (English)
Foundational world models must be both interactive and preserve spatiotemporal coherence for effective future planning with action choices. However, present models for long video generation have limited inherent world modeling capabilities due to two main challenges: compounding errors and insufficient memory mechanisms. We enhance image-to-video models with interactive capabilities through additional action conditioning and autoregressive framework, and reveal that compounding error is inherently irreducible in autoregressive video generation, while insufficient memory mechanism leads to incoherence of world models. We propose video retrieval augmented generation (VRAG) with explicit global state conditioning, which significantly reduces long-term compounding errors and increases spatiotemporal consistency of world models. In contrast, naive autoregressive generation with extended context windows and retrieval-augmented generation prove less effective for video generation, primarily due to the limited in-context learning capabilities of current video models. Our work illuminates the fundamental challenges in video world models and establishes a comprehensive benchmark for improving video generation models with internal world modeling capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。