arXiv:2410.07659cs.CV2024-10ICLR被引 7

用离散扩散模型生成高质量连贯视频,支持文本与草图驱动。

MotionAura: Generating High-Quality and Motion Consistent Videos using Discrete Diffusion

  • 采用向量量化扩散模型,将视频潜在空间离散化以捕捉复杂运动。
  • 在文本到视频生成中实现时序一致性,重构质量达当前最优水平。
  • 适合需要高保真视频生成与草图引导编辑的研究者与创作者。

视频数据的时空复杂性给压缩、生成和修复等任务带来巨大挑战。本文提出四项关键贡献:首先,引入3D Mobile Inverted Vector-Quantization VAE(3D-MBQ-VAE),结合变分自编码器与掩码标记建模,通过全帧掩码训练策略实现优异的时序一致性和最先进的重建质量;其次,提出MotionAura文本到视频生成框架,利用向量量化扩散模型离散潜在空间,捕捉复杂运动动态,生成与文本提示对齐的时序连贯视频;第三,设计基于频谱变换的去噪网络,通过傅里叶变换在频域处理视频数据,有效捕捉全局上下文与长程依赖,提升生成与去噪质量;最后,提出草图引导视频修复这一下游任务,采用低秩适配(LoRA)实现参数高效微调。模型在多个基准测试中达到最先进性能。本工作为时空建模与用户驱动视频内容操作提供稳健框架,代码、数据集与模型将开源。

原文摘要 · Abstract (English)

The spatio-temporal complexity of video data presents significant challenges in tasks such as compression, generation, and inpainting. We present four key contributions to address the challenges of spatiotemporal video processing. First, we introduce the 3D Mobile Inverted Vector-Quantization Variational Autoencoder (3D-MBQ-VAE), which combines Variational Autoencoders (VAEs) with masked token modeling to enhance spatiotemporal video compression. The model achieves superior temporal consistency and state-of-the-art (SOTA) reconstruction quality by employing a novel training strategy with full frame masking. Second, we present MotionAura, a text-to-video generation framework that utilizes vector-quantized diffusion models to discretize the latent space and capture complex motion dynamics, producing temporally coherent videos aligned with text prompts. Third, we propose a spectral transformer-based denoising network that processes video data in the frequency domain using the Fourier Transform. This method effectively captures global context and long-range dependencies for high-quality video generation and denoising. Lastly, we introduce a downstream task of Sketch Guided Video Inpainting. This task leverages Low-Rank Adaptation (LoRA) for parameter-efficient fine-tuning. Our models achieve SOTA performance on a range of benchmarks. Our work offers robust frameworks for spatiotemporal modeling and user-driven video content manipulation. We will release the code, datasets, and models in open-source.

视频生成扩散模型时序一致文本生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。