arXiv:2510.23561eess.IVcs.CV2025-10

改进视频编码中头部旋转的生成效果,降低40%~80%码率。

Revising Second Order Terms in Deep Animation Video Coding

  • 用全局旋转替代原模型的雅可比变换,提升头部转动处理能力。
  • 在包含头部旋转的视频上,P帧码率降低40%至80%。
  • 结合先进归一化技术稳定训练,提升生成视频视觉质量。

首阶运动模型(FOMM)是一种基于关键点信息生成人脸动画的生成模型,因其低码率和适中计算量,是视频通信的有前景方案。但其设计缺陷明显:通过图像扭曲生成动画,在强头部运动场景下表现不佳。本文聚焦头部旋转这一特定运动类型,提出用全局旋转替代原模型中的雅可比变换,显著提升对头部旋转视频的重建性能。实验表明,该优化使P帧码率降低40%至80%。同时,引入当前最先进的归一化技术于判别器,有效稳定对抗训练过程,保障生成视频的视觉质量。性能评估采用学习型度量指标LPIPS与DISTS,验证了优化的有效性。

原文摘要 · Abstract (English)

First Order Motion Model is a generative model that animates human heads based on very little motion information derived from keypoints. It is a promising solution for video communication because first it operates at very low bitrate and second its computational complexity is moderate compared to other learning based video codecs. However, it has strong limitations by design. Since it generates facial animations by warping source-images, it fails to recreate videos with strong head movements. This works concentrates on one specific kind of head movements, namely head rotations. We show that replacing the Jacobian transformations in FOMM by a global rotation helps the system to perform better on items with head-rotations while saving 40% to 80% of bitrate on P-frames. Moreover, we apply state-of-the-art normalization techniques to the discriminator to stabilize the adversarial training which is essential for generating visually appealing videos. We evaluate the performance by the learned metics LPIPS and DISTS to show the success our optimizations.

视频编码生成模型头部运动低码率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。