用少量关键帧生成完整视频,大幅降低存储成本。
Generative Latent Diffusion for Efficient Spatiotemporal Data Reduction
- 用变分自编码器+条件扩散模型压缩关键帧,实现高效生成重建
- 压缩比最高达10倍,相同误差下性能优于主流学习方法63%
- 适合需要低存储高保真视频还原的场景
生成模型在条件设定下表现出色,可视为一种数据压缩方式,其中条件作为紧凑表示。然而,其可控性有限且重建精度不足,限制了其在数据压缩中的实际应用。本文提出一种高效的潜空间扩散框架,通过将变分自编码器与条件扩散模型结合,仅将少量关键帧压缩至潜空间,并以此作为条件输入,通过生成插值重建其余帧,无需存储每帧的潜表示。该方法在保持高时空重建精度的同时显著降低存储开销。多个数据集上的实验表明,本方法压缩比最高可达规则类先进压缩器SZ3的10倍,相同重建误差下性能优于领先学习型方法63%。
原文摘要 · Abstract (English)
Generative models have demonstrated strong performance in conditional settings and can be viewed as a form of data compression, where the condition serves as a compact representation. However, their limited controllability and reconstruction accuracy restrict their practical application to data compression. In this work, we propose an efficient latent diffusion framework that bridges this gap by combining a variational autoencoder with a conditional diffusion model. Our method compresses only a small number of keyframes into latent space and uses them as conditioning inputs to reconstruct the remaining frames via generative interpolation, eliminating the need to store latent representations for every frame. This approach enables accurate spatiotemporal reconstruction while significantly reducing storage costs. Experimental results across multiple datasets show that our method achieves up to 10 times higher compression ratios than rule-based state-of-the-art compressors such as SZ3, and up to 63 percent better performance than leading learning-based methods under the same reconstruction error.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。