arXiv:2506.13419eess.IVcs.CV2025-06中稿 · ICMR2025被引 1

用音视频驱动压缩低码率人脸视频,提升同步与画质。

Audio-Visual Driven Compression for Low-Bitrate Talking Head Videos

  • 结合音频信号与3D运动特征,增强大角度头部动作建模。
  • 相比VVC降低22%码率,优于当前最先进方法8.5%。
  • 适合低带宽场景,显著改善唇语同步和面部细节还原。

人脸视频压缩在神经渲染与关键点方法推动下取得进展,但在低码率下仍面临大范围头部运动处理、唇音不同步及面部重建失真等挑战。为此,我们提出一种新型音视频驱动的视频编解码器,融合紧凑的3D运动特征与音频信号,有效建模大幅头部旋转并精准对齐唇部动作与语音,提升压缩效率与重建质量。在CelebV-HQ数据集上的实验表明,该方法相较VVC降低22%码率,比当前最先进的学习型编解码器降低8.5%。此外,在相近码率下,其唇音同步精度与视觉保真度均显著更优,展现出在带宽受限场景下的高效性。

原文摘要 · Abstract (English)

Talking head video compression has advanced with neural rendering and keypoint-based methods, but challenges remain, especially at low bit rates, including handling large head movements, suboptimal lip synchronization, and distorted facial reconstructions. To address these problems, we propose a novel audio-visual driven video codec that integrates compact 3D motion features and audio signals. This approach robustly models significant head rotations and aligns lip movements with speech, improving both compression efficiency and reconstruction quality. Experiments on the CelebV-HQ dataset show that our method reduces bitrate by 22% compared to VVC and by 8.5% over state-of-the-art learning-based codec. Furthermore, it provides superior lip-sync accuracy and visual fidelity at comparable bitrates, highlighting its effectiveness in bandwidth-constrained scenarios.

视频压缩音视频协同人脸生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。