arXiv:2410.05322cs.CV2024-10

用噪声调控让图像模型零样本生成连贯视频,兼顾细节与运动控制。

Noise Crystallization and Liquid Noise: Zero-shot Video Generation using Image Diffusion Models

  • 通过修改潜空间噪声实现图像模型零样本生成视频帧序列。
  • 噪声结晶化保持一致性但限于大运动,液态噪声提升灵活性无分辨率限制。
  • 适用于重光照、无缝放大和视频风格迁移,适合快速原型设计者。

尽管扩散模型在图像生成中表现强大,但生成一致且可控的视频仍是长期难题。视频模型需大量训练和计算资源,成本高且环境影响大,同时输出运动控制能力有限。本文提出一种新方法:通过增强图像扩散模型生成连续动画帧,无需训练任何视频参数(零样本),仅通过调整潜空间噪声即可实现。提出两种互补技术:噪声结晶化可保证一致性,但因潜嵌入尺寸减小而仅适用于大运动;液态噪声则牺牲一致性换取更大灵活性,无分辨率限制。核心思想还可拓展至重光照、无缝放大及视频风格迁移等应用。此外,研究了用于潜空间扩散模型的VAE嵌入,获得人可理解潜空间的理论洞见。

原文摘要 · Abstract (English)

Although powerful for image generation, consistent and controllable video is a longstanding problem for diffusion models. Video models require extensive training and computational resources, leading to high costs and large environmental impacts. Moreover, video models currently offer limited control of the output motion. This paper introduces a novel approach to video generation by augmenting image diffusion models to create sequential animation frames while maintaining fine detail. These techniques can be applied to existing image models without training any video parameters (zero-shot) by altering the input noise in a latent diffusion model. Two complementary methods are presented. Noise crystallization ensures consistency but is limited to large movements due to reduced latent embedding sizes. Liquid noise trades consistency for greater flexibility without resolution limitations. The core concepts also allow other applications such as relighting, seamless upscaling, and improved video style transfer. Furthermore, an exploration of the VAE embedding used for latent diffusion models is performed, resulting in interesting theoretical insights such as a method for human-interpretable latent spaces.

视频生成扩散模型零样本潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。