arXiv:2409.01199cs.CVeess.IV2024-09被引 38

提出可时空联合压缩视频的OD-VAE,提升扩散模型效率。

OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model

论文配图:OD-VAE: An Omni-dimensional Video Compressor for Improving Latent Video Diffusion Model
图 1 · 摘自论文原文
  • 用3D因果卷积实现视频时空联合压缩
  • 在保持高重建精度下压缩率显著提升
  • 支持任意长度视频生成,适合资源受限场景

变分自编码器(VAE)将视频压缩为潜在表示,是潜在视频扩散模型(LVDMs)的关键前置组件。在相同重建质量下,VAE对视频的压缩越充分,LVDMs的效率越高。然而,多数LVDMs采用二维图像VAE,仅在空间维度进行压缩,常忽略时间维度。如何在保证重建准确性的前提下,对视频进行有效的时间压缩仍缺乏研究。为此,本文提出一种全维压缩VAE——OD-VAE,可同时实现时空压缩。尽管更充分的压缩带来重建挑战,但通过精心设计仍能保持高重建精度。为平衡重建质量与压缩速度,提出了四种变体并进行了分析。此外,设计了一种新型尾部初始化方法以提高训练效率,并提出一种新推理策略,使OD-VAE可在有限显存下处理任意长度视频。在视频重建及基于LVDM的视频生成任务上的全面实验验证了方法的有效性与高效性。

原文摘要 · Abstract (English)

Variational Autoencoder (VAE), compressing videos into latent representations, is a crucial preceding component of Latent Video Diffusion Models (LVDMs). With the same reconstruction quality, the more sufficient the VAE's compression for videos is, the more efficient the LVDMs are. However, most LVDMs utilize 2D image VAE, whose compression for videos is only in the spatial dimension and often ignored in the temporal dimension. How to conduct temporal compression for videos in a VAE to obtain more concise latent representations while promising accurate reconstruction is seldom explored. To fill this gap, we propose an omni-dimension compression VAE, named OD-VAE, which can temporally and spatially compress videos. Although OD-VAE's more sufficient compression brings a great challenge to video reconstruction, it can still achieve high reconstructed accuracy by our fine design. To obtain a better trade-off between video reconstruction quality and compression speed, four variants of OD-VAE are introduced and analyzed. In addition, a novel tail initialization is designed to train OD-VAE more efficiently, and a novel inference strategy is proposed to enable OD-VAE to handle videos of arbitrary length with limited GPU memory. Comprehensive experiments on video reconstruction and LVDM-based video generation demonstrate the effectiveness and efficiency of our proposed methods.

视频压缩扩散模型3D卷积潜空间

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。