arXiv:2503.19604eess.IVcs.CV2025-03ICCV被引 11

用隐式扩散模型提升视频压缩效率,首次超越传统编码器。

GIViC: Generative Implicit Video Compression

  • 设计隐式扩散过程,分层渐进重构视频帧。
  • 在相同配置下,相比VVC、DCVC-FM等模型分别节省15.94%、22.46%码率。
  • 适合关注生成式压缩与高效视频编码的研究者。

尽管基于隐式神经表示(INRs)的视频压缩近年展现出巨大潜力,但在相同编码配置下,现有INR视频编码器仍未能达到传统或自编码器方法的最先进性能。为此,本文提出生成式隐式视频压缩框架GIViC,旨在突破此类方法的性能极限。GIViC受大型语言模型和扩散模型在建模长程依赖方面的启发,引入新型隐式扩散过程,实现从粗粒度全序列扩散到细粒度逐标记扩散的渐进采样。同时,集成一种新型分层门控线性注意力变压器(HGLA),沿尺度与时间轴双重分解全局依赖关系。在随机访问(RA)配置(YUV 4:2:0,GOPSize=32)下,GIViC相较于最新传统及神经编码器,在比特率上分别实现15.94%、22.46%和8.52%的BD-rate节省。据我们所知,GIViC是首个在该配置下超越VTM的INR基视频编码器。源代码将公开。

原文摘要 · Abstract (English)

While video compression based on implicit neural representations (INRs) has recently demonstrated great potential, existing INR-based video codecs still cannot achieve state-of-the-art (SOTA) performance compared to their conventional or autoencoder-based counterparts given the same coding configuration. In this context, we propose a Generative Implicit Video Compression framework, GIViC, aiming at advancing the performance limits of this type of coding methods. GIViC is inspired by the characteristics that INRs share with large language and diffusion models in exploiting long-term dependencies. Through the newly designed implicit diffusion process, GIViC performs diffusive sampling across coarse-to-fine spatiotemporal decompositions, gradually progressing from coarser-grained full-sequence diffusion to finer-grained per-token diffusion. A novel Hierarchical Gated Linear Attention-based transformer (HGLA), is also integrated into the framework, which dual-factorizes global dependency modeling along scale and sequential axes. The proposed GIViC model has been benchmarked against SOTA conventional and neural codecs using a Random Access (RA) configuration (YUV 4:2:0, GOPSize=32), and yields BD-rate savings of 15.94%, 22.46% and 8.52% over VVC VTM, DCVC-FM and NVRC, respectively. As far as we are aware, GIViC is the first INR-based video codec that outperforms VTM based on the RA coding configuration. The source code will be made available.

视频压缩隐式表示扩散模型生成式编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。