arXiv:2412.09600cs.CVcs.AI2024-12被引 17

用隐状态建模世界演化,生成更连贯的长视频。

Owl-1: Omni World Model for Consistent Long Video Generation

  • 用隐变量表示世界状态,动态更新以保持长期一致
  • 在VBench-I2V和VBench-Long上达到顶尖水平
  • 适合需要长时一致性视频生成的研究与应用

视频生成模型(VGMs)近年来受到广泛关注,被视为通用大视觉模型的有力候选。然而现有方法仅能单次生成短视频,通过迭代调用并以最后一帧为条件生成长视频,但最后一帧仅包含短期细粒度场景信息,导致长期不一致。为此,我们提出全知世界模型Owl-1,为长视频生成提供长期一致且全面的条件。由于视频是底层动态世界观测结果,我们提出在隐空间中建模长期演变,并用VGM将这些演变转化为视频。具体地,用一个隐状态变量表示世界,可解码为显式视频观测,这些观测用于预测时间动态,进而更新状态变量。动态演化与持久状态间的交互增强了视频的多样性和一致性。大量实验表明,Owl-1在VBench-I2V和VBench-Long上表现媲美最先进方法,验证了其生成高质量视频观测的能力。代码见:https://github.com/huang-yh/Owl。

原文摘要 · Abstract (English)

Video generation models (VGMs) have received extensive attention recently and serve as promising candidates for general-purpose large vision models. While they can only generate short videos each time, existing methods achieve long video generation by iteratively calling the VGMs, using the last-frame output as the condition for the next-round generation. However, the last frame only contains short-term fine-grained information about the scene, resulting in inconsistency in the long horizon. To address this, we propose an Omni World modeL (Owl-1) to produce long-term coherent and comprehensive conditions for consistent long video generation. As videos are observations of the underlying evolving world, we propose to model the long-term developments in a latent space and use VGMs to film them into videos. Specifically, we represent the world with a latent state variable which can be decoded into explicit video observations. These observations serve as a basis for anticipating temporal dynamics which in turn update the state variable. The interaction between evolving dynamics and persistent state enhances the diversity and consistency of the long videos. Extensive experiments show that Owl-1 achieves comparable performance with SOTA methods on VBench-I2V and VBench-Long, validating its ability to generate high-quality video observations. Code: https://github.com/huang-yh/Owl.

视频生成长视频一致性隐状态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。