无需训练即可从文本生成无缝循环视频,支持任意时长循环。
Mobius: Text to Seamless Looping Video Generation via Latent Shift
- 通过潜空间循环与渐进移位实现多帧去噪,保持时间一致性。
- 生成视频可任意长度循环,超越模型原始上下文限制。
- 不依赖图像输入,动态更丰富,视觉质量更高。
我们提出 Mobius,一种直接从文本描述生成无缝循环视频的新方法,无需用户标注或额外训练。该方法复用预训练的视频潜空间扩散模型,在推理阶段通过连接视频起始与终止噪声构建潜空间循环,并逐步将首帧潜变量移向末帧,在每一步中改变去噪上下文的同时维持全程一致性。潜空间循环长度可任意设定,突破了原视频扩散模型上下文长度限制。相比以往的动态照片(cinemagraphs),本方法无需初始图像作为外观参考,从而避免运动受限问题,可生成更具动态性与高质量的视频。我们通过多组实验验证了方法的有效性,涵盖多种场景。所有代码将公开。
原文摘要 · Abstract (English)
We present Mobius, a novel method to generate seamlessly looping videos from text descriptions directly without any user annotations, thereby creating new visual materials for the multi-media presentation. Our method repurposes the pre-trained video latent diffusion model for generating looping videos from text prompts without any training. During inference, we first construct a latent cycle by connecting the starting and ending noise of the videos. Given that the temporal consistency can be maintained by the context of the video diffusion model, we perform multi-frame latent denoising by gradually shifting the first-frame latent to the end in each step. As a result, the denoising context varies in each step while maintaining consistency throughout the inference process. Moreover, the latent cycle in our method can be of any length. This extends our latent-shifting approach to generate seamless looping videos beyond the scope of the video diffusion model's context. Unlike previous cinemagraphs, the proposed method does not require an image as appearance, which will restrict the motions of the generated results. Instead, our method can produce more dynamic motion and better visual quality. We conduct multiple experiments and comparisons to verify the effectiveness of the proposed method, demonstrating its efficacy in different scenarios. All the code will be made available.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。