arXiv:2502.17085cs.CVeess.IV2025-02被引 6

用智能带宽生成面部视频,实现全码率范围的高质量压缩。

Pleno-Generation: A Scalable Generative Face Video Compression Framework with Bandwidth Intelligence

  • 通过可扩展表示与分层重建,让压缩码流具备不同粒度的智能性。
  • 在更广码率范围内实现高保真面部视频重建,性能优于现有生成式编码。
  • 适合需要全码率覆盖和高画质的实时视频通信应用。

基于生成模型的紧凑视频压缩通常局限于较窄的码率范围,且多聚焦于超低码率场景。业界日益达成共识:生成编码应支持全码率覆盖。然而,这极具挑战性,因生成与压缩目标和权衡本质不同。所提出的Pleno-Generation(PGen)框架通过带宽智能机制,在更宽的带宽范围内进行生成,显著提升了视频编码的鲁棒性。本研究以人脸视频压缩为切入点,提出一种范式转变:优先保障高保真重建而非追求极小码流。PGen采用可扩展表示与分层重建,实现生成式人脸视频压缩(GFVC),使码流具备多层次智能。实验表明,PGen可显著提升现有GFVC算法的高保真与真实感表现;同时提供更大灵活性,在多种质量评估下展现更优率失真性能,覆盖更广码率区间。相比最新通用视频编码(VVC)标准,该方案在感知级评价中达到相当的Bjøntegaard-delta-rate节省效果。

原文摘要 · Abstract (English)

Generative model based compact video compression is typically operated within a relative narrow range of bitrates, and often with an emphasis on ultra-low rate applications. There has been an increasing consensus in the video communication industry that full bitrate coverage should be enabled by generative coding. However, this is an extremely difficult task, largely because generation and compression, although related, have distinct goals and trade-offs. The proposed Pleno-Generation (PGen) framework distinguishes itself through its exceptional capabilities in ensuring the robustness of video coding by utilizing a wider range of bandwidth for generation via bandwidth intelligence. In particular, we initiate our research of PGen with face video coding, and PGen offers a paradigm shift that prioritizes high-fidelity reconstruction over pursuing compact bitstream. The novel PGen framework leverages scalable representation and layered reconstruction for Generative Face Video Compression (GFVC), in an attempt to imbue the bitstream with intelligence in different granularity. Experimental results illustrate that the proposed PGen framework can facilitate existing GFVC algorithms to better deliver high-fidelity and faithful face videos. In addition, the proposed framework can allow a greater space of flexibility for coding applications and show superior RD performance with a much wider bitrate range in terms of various quality evaluations. Moreover, in comparison with the latest Versatile Video Coding (VVC) codec, the proposed scheme achieves competitive Bjøntegaard-delta-rate savings for perceptual-level evaluations.

生成压缩人脸视频带宽智能率失真优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。