arXiv:2410.15105cs.CV2024-10被引 6

用补充增强信息提升人脸视频生成压缩质量,已进入国际标准制定阶段。

Standardizing Generative Face Video Compression using Supplemental Enhancement Information

  • 通过嵌入关键点等紧凑表示,利用SEI实现生成式人脸视频压缩。
  • 相比最新VVC标准,在相同码率下显著提升重建质量。
  • 支持用户定制动画与元宇宙应用,适合未来生成式视频系统使用。

本文提出一种基于补充增强信息(SEI)的生成式人脸视频压缩(GFVC)方法,将人脸视频信号的紧凑空间与时间表示(如2D/3D关键点、面部语义和紧凑特征)编码为SEI消息并插入码流。目前,该基于SEI的GFVC方案已被国际电信联盟(ITU-T)与ISO/IEC JTC 1/SC 29联合视频专家组(JVET)纳入《通用补充增强信息》(VSEI)标准草案修订案,即将成为ITU-T H.274 | ISO/IEC 23002-7的新版本。据作者所知,这是首个针对生成式视频压缩的标准化工作。该方法不仅借助先进生成技术显著提升了早期模型基编码(MBC)的重建质量,还为未来GFVC应用建立了新的SEI定义。实验表明,该方案在率失真性能上显著优于最新版通用视频编码(VVC)标准,同时具备支持用户自定义动画/滤镜及元宇宙相关应用的潜力。

原文摘要 · Abstract (English)

This paper proposes a Generative Face Video Compression (GFVC) approach using Supplemental Enhancement Information (SEI), where a series of compact spatial and temporal representations of a face video signal (e.g., 2D/3D keypoints, facial semantics and compact features) can be coded using SEI messages and inserted into the coded video bitstream. At the time of writing, the proposed GFVC approach using SEI messages has been included into a draft amendment of the Versatile Supplemental Enhancement Information (VSEI) standard by the Joint Video Experts Team (JVET) of ISO/IEC JTC 1/SC 29 and ITU-T SG21, which will be standardized as a new version of ITU-T H.274 | ISO/IEC 23002-7. To the best of the authors' knowledge, the JVET work on the proposed SEI-based GFVC approach is the first standardization activity for generative video compression. The proposed SEI approach has not only advanced the reconstruction quality of early-day Model-Based Coding (MBC) via the state-of-the-art generative technique, but also established a new SEI definition for future GFVC applications and deployment. Experimental results illustrate that the proposed SEI-based GFVC approach can achieve remarkable rate-distortion performance compared with the latest Versatile Video Coding (VVC) standard, whilst also potentially enabling a wide variety of functionalities including user-specified animation/filtering and metaverse-related applications.

生成式压缩人脸视频SEI标准制定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。