用生成模型实现超低码率高保真人脸视频压缩,超越现有标准。
Generative Models at the Frontier of Compression: A Survey on Generative Face Video Coding
- 通过语义感知的隐变量编码,实现人脸动态的紧凑表示。
- 在极低码率下重建高保真人脸视频,性能远超最新VVC标准。
- 构建首个面向主观评价的大规模数据集,推动标准制定与应用落地。
深度生成模型的兴起极大推进了视频压缩技术,重塑了人脸视频编码范式,凭借其强大的语义感知表征与逼真合成能力。生成式人脸视频编码(GFVC)处于这一变革前沿,可在编码端将复杂人脸动态压缩为紧凑的隐变量码流,并在解码端利用强大生成模型从压缩隐码中重建高质量人脸信号。该设计使超低码率下实现高保真人脸视频通信成为可能,显著超越最新通用视频编码(VVC)标准。为推动基础研究并加速GFVC发展,本文首次系统综述该领域技术,梳理不同特征表示与优化策略,开展全面基准测试;构建包含主观评分(MOS)的大规模GFVC压缩人脸视频数据库,以识别适配该场景的质量评估指标;总结统一语法架构下的标准化潜力,并提出低复杂度系统原型,助力未来实际部署。最后,展望工业应用场景,分析当前挑战与未来机遇。
原文摘要 · Abstract (English)
The rise of deep generative models has greatly advanced video compression, reshaping the paradigm of face video coding through their powerful capability for semantic-aware representation and lifelike synthesis. Generative Face Video Coding (GFVC) stands at the forefront of this revolution, which could characterize complex facial dynamics into compact latent codes for bitstream compactness at the encoder side and leverages powerful deep generative models to reconstruct high-fidelity face signal from the compressed latent codes at the decoder side. As such, this well-designed GFVC paradigm could enable high-fidelity face video communication at ultra-low bitrate ranges, far surpassing the capabilities of the latest Versatile Video Coding (VVC) standard. To pioneer foundational research and accelerate the evolution of GFVC, this paper presents the first comprehensive survey of GFVC technologies, systematically bridging critical gaps between theoretical innovation and industrial standardization. In particular, we first review a broad range of existing GFVC methods with different feature representations and optimization strategies, and conduct a thorough benchmarking analysis. In addition, we construct a large-scale GFVC-compressed face video database with subjective Mean Opinion Scores (MOSs) based on human perception, aiming to identify the most appropriate quality metrics tailored to GFVC. Moreover, we summarize the GFVC standardization potentials with a unified high-level syntax and develop a low-complexity GFVC system which are both expected to push forward future practical deployments and applications. Finally, we envision the potential of GFVC in industrial applications and deliberate on the current challenges and future opportunities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。