通过分层参考结构提升生成式视频压缩的效率与画质。
Generative Video Compression Based on Hierarchical Referencing

- 构建分层参考与质量结构,优先保护高质帧
- 在多个基准上实现LPIPS和DISTS指标50.5%和54.0%的码率降低
- 适合关注生成式视频压缩与扩散模型应用的研究者
基于扩散模型的生成式视频压缩在提升主观画质方面展现出巨大潜力,但现有方法在潜在帧编码时未合理设计参考与质量结构,且忽视帧级质量差异对去噪过程的影响,限制了编码效率并加剧了重建中的伪影传播。本文提出GVCHR(基于分层参考的生成式视频压缩),其核心思想是将潜在帧分层组织,使高质参考帧同时优化编码与重建。在编码端,将分层参考结构与分层质量结构结合,为更频繁被用作参考的低层帧分配更多比特;在此基础上引入分层时间上下文挖掘,有效利用短时与长时时间上下文。在重建端,将编码端的层次结构融入视频扩散变压器的分层注意力适配器中,通过分层注意力机制限制每帧仅能访问同层或更低层参考帧,从而减少去噪过程中的伪影传播。实验表明,与先前最优方法相比,GVCHR在多个基准上分别实现了LPIPS和DISTS指标50.5%和54.0%的BD-rate提升,同时显著改善视觉质量。
原文摘要 · Abstract (English)
Diffusion-based generative video compression has emerged as a promising paradigm to improve perceptual quality, where latent frames are required to be encoded efficiently while serving as denoising conditions. However, existing methods neither carefully design reference and quality structures during latent coding nor account for the impact of frame-level quality variation on denoising procedure, which limits coding efficiency and aggravates artifact propagation during generative reconstruction. In this paper, we propose GVCHR, Generative Video Compression based on Hierarchical Referencing. The key idea is to organize latent frames hierarchically, where the selected high-quality references benefit both latent coding and generative reconstruction. In latent coding, GVCHR couples a hierarchical reference structure with a hierarchical quality structure, assigning more bits to lower-layer frames that are reused more frequently as references. Built on this design, we introduce Hierarchical Temporal Context Mining to exploits complementary short- and long-term temporal context for effective latent coding. In generative reconstruction, the coding-side hierarchy is incorporated into a Hierarchical Attentive Adapter which is attached to a video diffusion transformer. This adapter uses hierarchical attention to restrict each latent frame to attend only to the same- or lower-layer references, thereby reducing artifact propagation during denoising. Experiments validate GVCHR on multiple benchmarks. Compared with the previous state-of-the-art method, GVCHR achieves 50.5% and 54.0% BD-rate gains in terms of LPIPS and DISTS, respectively, while also delivering clearly improved visual quality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。