通过去除视频隐变量高频成分实现压缩,提升生成效果
Latent-Compressed Variational Autoencoder for Video Diffusion Models

- 不直接减少通道数,而是滤除隐变量高频信息
- 相同压缩比下,重建质量优于强基线方法
- 适合追求高质量视频生成的扩散模型研究者
在潜空间扩散模型中,视频变分自编码器(VAE)通常需要足够多的潜通道以保证高质量视频重建。然而近期研究表明,过多的潜通道会阻碍潜空间扩散模型的收敛,并降低其生成性能,即使重建质量仍较高。本文提出一种潜压缩方法,通过去除视频潜表示中的高频成分,而非直接减少通道数,从而避免重建保真度下降。实验结果表明,该方法在保持相同整体压缩比的情况下,相比强基线实现了更优的视频重建质量。
原文摘要 · Abstract (English)
Video variational autoencoders (VAEs) used in latent diffusion models typically require a sufficiently large number of latent channels to ensure high-quality video reconstruction. However, recent studies have revealed that an excessive number of latent channels can impede the convergence of latent diffusion models and deteriorate their generative performance, even when reconstruction quality remains high. We propose a latent compression method that removes high-frequency components in video latent representations rather than directly reducing the number of channels, which often compromises reconstruction fidelity. Experimental results demonstrate that the proposed method achieves superior video reconstruction quality compared to strong baselines while maintaining the same overall compression ratio.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。