arXiv:2503.09787eess.IVcs.CV2025-03中稿 · DCC2025被引 3

用前后关键帧双向编码,大幅降低说话头视频码率。

Bidirectional Learned Facial Animation Codec for Low Bitrate Talking Head Videos

  • 结合前后关键帧,动态选择参考帧提升重建质量
  • 相比最新编码标准,码率降低35%以上
  • 适合低码率下高质量说话头视频生成

现有深度面部动画编码技术通过生成模型压缩说话头视频,仅编码关键帧和非关键帧的特征点,再由单个关键帧和特征点重建目标帧。然而这类单向方法依赖单一关键帧,难以准确捕捉大范围头部运动,导致面部区域失真。本文提出一种新型双向学习动画编解码器,利用过去和未来关键帧生成自然面部视频。首先,在双向参考引导辅助流增强(BRG-ASE)过程中,为非关键帧引入紧凑辅助流,并自适应选择过去或未来关键帧进行增强,轻微增加码率但显著提升画质;其次,在双向参考引导视频重建(BRG-VRec)过程中,使用选定的关键帧与辅助帧联合重建目标帧。大量实验表明,该方法相比最新基于动画的视频编码器减少55%码率,相比最新视频编码标准VVC减少35%码率,在说话头视频数据集上实现高效画质提升。

原文摘要 · Abstract (English)

Existing deep facial animation coding techniques efficiently compress talking head videos by applying deep generative models. Instead of compressing the entire video sequence, these methods focus on compressing only the keyframe and the keypoints of non-keyframes (target frames). The target frames are then reconstructed by utilizing a single keyframe, and the keypoints of the target frame. Although these unidirectional methods can reduce the bitrate, they rely on a single keyframe and often struggle to capture large head movements accurately, resulting in distortions in the facial region. In this paper, we propose a novel bidirectional learned animation codec that generates natural facial videos using past and future keyframes. First, in the Bidirectional Reference-Guided Auxiliary Stream Enhancement (BRG-ASE) process, we introduce a compact auxiliary stream for non-keyframes, which is enhanced by adaptively selecting one of two keyframes (past and future). This stream improves video quality with a slight increase in bitrate. Then, in the Bidirectional Reference-Guided Video Reconstruction (BRG-VRec) process, we animate the adaptively selected keyframe and reconstruct the target frame using both the animated keyframe and the auxiliary frame. Extensive experiments demonstrate a 55% bitrate reduction compared to the latest animation based video codec, and a 35% bitrate reduction compared to the latest video coding standard, Versatile Video Coding (VVC) on a talking head video dataset. It showcases the efficiency of our approach in improving video quality while simultaneously decreasing bitrate.

视频编码面部动画低码率双向建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。