用隐式运动变换提升人体视频生成编码的压缩与重建质量
Rethinking Generative Human Video Coding with Implicit Motion Transformation
- 将人体动作转为隐式运动信号,替代显式运动场
- 在人体视频上实现更高压缩率与更清晰重建
- 适合关注视频生成与高效编码的研究者
传统混合视频编码之外,生成式视频编码通过将高维信号转换为紧凑特征表示,实现编码端比特流紧凑,并在解码端利用显式运动场作为中间监督以实现高质量重建,已在人脸视频压缩中取得显著成功。然而,相比人脸视频,人体视频因运动模式更复杂多样,使用显式运动引导进行生成式人体视频编码(GHVC)时,重建结果易出现严重失真和运动不准。本文指出显式运动方法在人体视频压缩中的局限性,提出隐式运动变换(IMT)以改进GHVC性能。具体而言,将复杂人体信号压缩为紧凑视觉特征,并将其转换为隐式运动引导用于信号重建。实验表明,所提IMT范式可有效提升GHVC的压缩效率与重建保真度。
原文摘要 · Abstract (English)
Beyond traditional hybrid-based video codec, generative video codec could achieve promising compression performance by evolving high-dimensional signals into compact feature representations for bitstream compactness at the encoder side and developing explicit motion fields as intermediate supervision for high-quality reconstruction at the decoder side. This paradigm has achieved significant success in face video compression. However, compared to facial videos, human body videos pose greater challenges due to their more complex and diverse motion patterns, i.e., when using explicit motion guidance for Generative Human Video Coding (GHVC), the reconstruction results could suffer severe distortions and inaccurate motion. As such, this paper highlights the limitations of explicit motion-based approaches for human body video compression and investigates the GHVC performance improvement with the aid of Implicit Motion Transformation, namely IMT. In particular, we propose to characterize complex human body signal into compact visual features and transform these features into implicit motion guidance for signal reconstruction. Experimental results demonstrate the effectiveness of the proposed IMT paradigm, which can facilitate GHVC to achieve high-efficiency compression and high-fidelity synthesis.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。