arXiv:2602.01340cs.CV2026-02

让VAE支持多级时间压缩,提升视频生成效率

MTC-VAE: Multi-Level Temporal Compression with Content Awareness

  • 将固定压缩率VAE改造为支持多级时间压缩
  • 在高压缩率下保持性能,避免效率下降
  • 适合需要高效视频生成的扩散模型研究者

潜在视频扩散模型(LVDMs)依赖变分自编码器(VAEs)将视频压缩为紧凑的潜在表示。对于连续变分自编码器(VAEs),更高的压缩率更受青睐;然而,若不扩展隐藏通道维度,增加采样层会导致效率显著下降。本文提出一种技术,将固定压缩率的VAE转换为支持多级时间压缩的模型,提供一种简单且最小化微调的方法,以缓解高压缩率下的性能下降问题。此外,我们研究了不同压缩级别对具有不同特征视频片段的影响,提供了所提方法有效性的实证证据。还探讨了该多级时间压缩VAE与基于扩散的生成模型DiT的集成,展示了其在这些框架中的成功协同训练与兼容性,揭示了多级时间压缩的潜在应用价值。

原文摘要 · Abstract (English)

Latent Video Diffusion Models (LVDMs) rely on Variational Autoencoders (VAEs) to compress videos into compact latent representations. For continuous Variational Autoencoders (VAEs), achieving higher compression rates is desirable; yet, the efficiency notably declines when extra sampling layers are added without expanding the dimensions of hidden channels. In this paper, we present a technique to convert fixed compression rate VAEs into models that support multi-level temporal compression, providing a straightforward and minimal fine-tuning approach to counteract performance decline at elevated compression rates.Moreover, we examine how varying compression levels impact model performance over video segments with diverse characteristics, offering empirical evidence on the effectiveness of our proposed approach. We also investigate the integration of our multi-level temporal compression VAE with diffusion-based generative models, DiT, highlighting successful concurrent training and compatibility within these frameworks. This investigation illustrates the potential uses of multi-level temporal compression.

视频生成扩散模型VAE压缩

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。