arXiv:2410.08485eess.IVcs.CV2024-10被引 8

提出PFVC框架,用自适应视觉令牌实现更稳定高效的面部视频压缩。

Beyond GFVC: A Progressive Face Video Compression Framework with Adaptive Visual Tokens

  • 分阶段将人脸信号转为自适应视觉令牌,提升编码灵活性。
  • 在低码率下仍保持高质量重建,优于最新VVC和GFVC算法。
  • 适合对压缩效率与画质平衡有高要求的应用场景。

近年来,深度生成模型显著推动了面部视频编码的发展,实现了优异的码率-失真性能和多样应用功能。不同于传统混合编码范式,基于深度生成模型和早期模型基编码(MBC)理念的生成式面部视频压缩(GFVC)可实现视觉人脸信号的紧凑表示与真实重建,从而达成超低码率面部视频通信。然而,现有GFVC算法常面临重建质量不稳定和码率范围有限的问题。为此,本文提出一种新型渐进式面部视频压缩框架PFVC,通过自适应视觉令牌实现重建鲁棒性与带宽智能性的卓越权衡。具体而言,所提PFVC的编码器以渐进方式将高维人脸信号映射为自适应视觉令牌,解码器则可按不同粒度层级重构这些令牌以进行运动估计与信号合成。实验结果表明,相比最新通用视频编码(VVC)标准及当前最优的GFVC算法,所提PFVC框架在编码灵活性和码率-失真性能上均表现更优。

原文摘要 · Abstract (English)

Recently, deep generative models have greatly advanced the progress of face video coding towards promising rate-distortion performance and diverse application functionalities. Beyond traditional hybrid video coding paradigms, Generative Face Video Compression (GFVC) relying on the strong capabilities of deep generative models and the philosophy of early Model-Based Coding (MBC) can facilitate the compact representation and realistic reconstruction of visual face signal, thus achieving ultra-low bitrate face video communication. However, these GFVC algorithms are sometimes faced with unstable reconstruction quality and limited bitrate ranges. To address these problems, this paper proposes a novel Progressive Face Video Compression framework, namely PFVC, that utilizes adaptive visual tokens to realize exceptional trade-offs between reconstruction robustness and bandwidth intelligence. In particular, the encoder of the proposed PFVC projects the high-dimensional face signal into adaptive visual tokens in a progressive manner, whilst the decoder can further reconstruct these adaptive visual tokens for motion estimation and signal synthesis with different granularity levels. Experimental results demonstrate that the proposed PFVC framework can achieve better coding flexibility and superior rate-distortion performance in comparison with the latest Versatile Video Coding (VVC) codec and the state-of-the-art GFVC algorithms. The project page can be found at https://github.com/Berlin0610/PFVC.

视频压缩生成模型自适应令牌人脸编码

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。