arXiv:2501.13528cs.CVcs.LG2025-01被引 29

用扩散模型提升视频压缩感知质量,加速推理并支持可变码率。

Diffusion-based Perceptual Neural Video Compression with Temporal Diffusion Information Reuse

  • 结合时序上下文与当前帧重构,引导扩散模型生成高质量视频。
  • 引入时序信息复用策略,推理效率显著提升且质量损失小。
  • 通过量化参数调制特征,实现鲁棒的可变码率压缩,适合实际应用。

近期,基础扩散模型在图像压缩中备受关注,但在视频压缩中的应用仍鲜有探索。本文提出 DiffVC,一种基于扩散模型的感知神经视频压缩框架,有效融合基础扩散模型与视频条件编码范式。该框架利用先前解码帧的时序上下文及当前帧的重构隐空间表示,引导扩散模型生成高质量结果。为加速扩散模型的迭代推理过程,提出时序扩散信息复用(TDIR)策略,通过复用前帧扩散信息,在极小性能损失下显著提升推理效率。此外,针对不同码率下的失真差异问题,提出基于量化参数提示(QPP)机制,将量化参数作为提示输入基础扩散模型,显式调控中间特征,从而构建鲁棒的可变码率扩散神经压缩框架。实验表明,所提方法在感知指标与视觉质量方面均表现优异。

原文摘要 · Abstract (English)

Recently, foundational diffusion models have attracted considerable attention in image compression tasks, whereas their application to video compression remains largely unexplored. In this article, we introduce DiffVC, a diffusion-based perceptual neural video compression framework that effectively integrates foundational diffusion model with the video conditional coding paradigm. This framework uses temporal context from previously decoded frame and the reconstructed latent representation of the current frame to guide the diffusion model in generating high-quality results. To accelerate the iterative inference process of diffusion model, we propose the Temporal Diffusion Information Reuse (TDIR) strategy, which significantly enhances inference efficiency with minimal performance loss by reusing the diffusion information from previous frames. Additionally, to address the challenges posed by distortion differences across various bitrates, we propose the Quantization Parameter-based Prompting (QPP) mechanism, which utilizes quantization parameters as prompts fed into the foundational diffusion model to explicitly modulate intermediate features, thereby enabling a robust variable bitrate diffusion-based neural compression framework. Experimental results demonstrate that our proposed solution delivers excellent performance in both perception metrics and visual quality.

视频压缩扩散模型感知质量可变码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。