arXiv:2412.04452cs.CV2024-12

四平面分解压缩视频潜在空间,加速生成模型训练与推理。

Factorized Video Autoencoders for Efficient Generative Modelling

  • 将视频数据映射到四平面因子化潜在空间,压缩率高且尺寸增长慢。
  • 在保持高质量重建的同时,生成速度提升显著,内存占用大幅降低。
  • 适用于条件生成任务,如类别生成、帧预测和视频插值。

潜在变量生成模型已成为图像与视频合成等生成任务的强大工具。这类模型依赖预训练的自编码器将高分辨率数据映射至压缩后的低维潜在空间,从而以更低计算成本构建生成模型。尽管有效,直接将潜在变量模型应用于视频等高维领域仍面临训练与推理效率挑战。本文提出一种自编码器,将体数据投影到四平面因子化潜在空间,其维度随输入规模呈亚线性增长,特别适合视频等高维数据。该设计可轻松集成至多种条件生成任务中的潜在扩散模型(LDM),如类别条件生成、帧预测与视频插值。结果表明,所提四平面潜在空间在实现高压缩比的同时,仍能保留高质量重建所需的丰富表征,同时使LDM在速度与内存上获得显著提升。

原文摘要 · Abstract (English)

Latent variable generative models have emerged as powerful tools for generative tasks including image and video synthesis. These models are enabled by pretrained autoencoders that map high resolution data into a compressed lower dimensional latent space, where the generative models can subsequently be developed while requiring fewer computational resources. Despite their effectiveness, the direct application of latent variable models to higher dimensional domains such as videos continues to pose challenges for efficient training and inference. In this paper, we propose an autoencoder that projects volumetric data onto a four-plane factorized latent space that grows sublinearly with the input size, making it ideal for higher dimensional data like videos. The design of our factorized model supports straightforward adoption in a number of conditional generation tasks with latent diffusion models (LDMs), such as class-conditional generation, frame prediction, and video interpolation. Our results show that the proposed four-plane latent space retains a rich representation needed for high-fidelity reconstructions despite the heavy compression, while simultaneously enabling LDMs to operate with significant improvements in speed and memory.

视频生成潜在空间自编码器扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。