arXiv:2504.10567cs.CVeess.IV2025-04被引 13

高压缩高速度高质量视频自编码器,支持移动端实时解码

H3AE: High Compression, High Speed, and High Quality AutoEncoder for Video Diffusion Models

  • 设计高效架构并优化计算分布,实现超高压缩比
  • 提出新损失函数,重建质量显著优于现有方法
  • 单模型支持多任务,适合视频生成与移动部署

自编码器(AE)是图像与视频生成中潜在扩散模型成功的关键,可降低去噪分辨率并提升效率。然而,网络设计、压缩率和训练策略方面仍存在未充分探索的空间。本文系统分析架构选择,优化计算分布,构建一系列高效且高压缩的视频自编码器,可在GPU和移动设备上实现实时解码。提出统一的全训练目标,使普通自编码器与图像条件的I2V VAE在单一VAE网络中实现多功能性,并提升质量。此外,提出新颖的潜在一致性损失,相比LPIPS、GAN和DWT等先前辅助损失,在质量提升与简洁性上均表现更优。H3AE实现超高压缩比与实时解码速度,在重建指标上大幅超越已有方法。最终通过在潜在空间训练DiT,验证其支持快速、高质量文本到视频生成的能力。

原文摘要 · Abstract (English)

Autoencoder (AE) is the key to the success of latent diffusion models for image and video generation, reducing the denoising resolution and improving efficiency. However, the power of AE has long been underexplored in terms of network design, compression ratio, and training strategy. In this work, we systematically examine the architecture design choices and optimize the computation distribution to obtain a series of efficient and high-compression video AEs that can decode in real time even on mobile devices. We also propose an omni-training objective to unify the design of plain Autoencoder and image-conditioned I2V VAE, achieving multifunctionality in a single VAE network but with enhanced quality. In addition, we propose a novel latent consistency loss that provides stable improvements in reconstruction quality. Latent consistency loss outperforms prior auxiliary losses including LPIPS, GAN and DWT in terms of both quality improvements and simplicity. H3AE achieves ultra-high compression ratios and real-time decoding speed on GPU and mobile, and outperforms prior arts in terms of reconstruction metrics by a large margin. We finally validate our AE by training a DiT on its latent space and demonstrate fast, high-quality text-to-video generation capability.

视频生成自编码器扩散模型移动部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。