arXiv:2606.22371eess.IVcs.CV2026-06被引 1

零样本视频压缩框架,用预训练扩散模型实现低延迟高清重建。

ZeroGVC: Zero-Shot Generative Video Compression with Autoregressive Diffusion Priors

论文配图:ZeroGVC: Zero-Shot Generative Video Compression with Autoregressive Diffusion Priors
图 1 · 摘自论文原文
  • 用自回归扩散先验直接生成视频帧,无需额外训练。
  • 在极低码率下实现超越现有方法的视觉质量。
  • 适合追求极致压缩效率与画质的视频系统开发者。

近期生成式视频压缩方法利用强大的生成先验实现高质量重构,但多数需额外训练以适配紧凑表示。本文提出零样本生成视频压缩框架 ZeroGVC,利用预训练的自回归扩散先验实现低延迟视频重建。ZeroGVC 使用图像编码器对每组图像(GOP)的第一帧进行编码,并通过基于码本引导的自回归潜在压缩表示后续P帧。该设计源于观察:去噪扩散码本模型在少步一致性采样中表现优异。通过选择可复现的码本噪声向量组合,ZeroGVC 引导潜在空间去噪轨迹至目标P帧,同时确保解码器仅用少数去噪步骤即可复现相同轨迹。此外,设计了可选的双向参考模式,利用下一I帧上下文缓解误差传播,且不增加比特率开销。在标准视频压缩基准上的大量实验表明,ZeroGVC 在无额外训练的前提下,于超低码率下实现了更优的感知重建质量。

原文摘要 · Abstract (English)

Recent generative video compression methods leverage powerful generative priors to achieve perceptually pleasing reconstructions. However, most existing approaches require additional training to adapt generative models to produce realistic reconstructions from compact representations. In this paper, we propose ZeroGVC, a zero-shot generative video compression framework that leverages pretrained autoregressive diffusion priors for low-delay video reconstruction. ZeroGVC encodes the first frame of each group of pictures (GOP) with an image codec and represents subsequent P-frames through Codebook-Guided Autoregressive Latent Compression. This design is motivated by our observation that the compression scheme of denoising diffusion codebook models is effective in few-step consistency sampling. By selecting compact combinations of reproducible codebook noise vectors, ZeroGVC steers the latent denoising trajectory toward the target P-frame while allowing the decoder to reproduce the same trajectory in only a few denoising steps. In addition, we design an optional bidirectional reference mode that mitigates error propagation by leveraging the next I-frame context without introducing any additional bitrate overhead. Extensive experiments on standard video compression benchmarks demonstrate that ZeroGVC achieves superior perceptual reconstruction quality at ultra-low bitrates without any additional training.

视频压缩生成模型扩散模型零样本

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。