arXiv:2507.20452cs.CV2025-07

联合学习3D人脸与说话头模型,实现更自然的口型同步。

JOLT3D: Joint Learning of Talking Heads and 3DMM Parameters with Application to Lip-Sync

  • 联合优化3DMM与说话头生成,提升面部质量
  • 通过混合形状表示精准控制嘴部区域,减少口型抖动
  • 适合需要高质量口型同步的数字人、虚拟主播应用

本文重新审视3DMM在说话头合成中的有效性,提出联合学习3D人脸重建模型与说话头生成模型的方法。该方法获得针对说话头合成优化的基于FACS的混合形状表示,突破了以往仅拟合2D关键点或依赖预训练重建模型的局限。不仅提升了生成人脸的质量,还利用混合形状表示实现仅修改嘴部区域的音频驱动口型同步。为此,我们设计了一种新型口型同步流程,将原始下巴轮廓与口型同步后的下巴轮廓解耦,显著减少嘴部附近的闪烁现象。

原文摘要 · Abstract (English)

In this work, we revisit the effectiveness of 3DMM for talking head synthesis by jointly learning a 3D face reconstruction model and a talking head synthesis model. This enables us to obtain a FACS-based blendshape representation of facial expressions that is optimized for talking head synthesis. This contrasts with previous methods that either fit 3DMM parameters to 2D landmarks or rely on pretrained face reconstruction models. Not only does our approach increase the quality of the generated face, but it also allows us to take advantage of the blendshape representation to modify just the mouth region for the purpose of audio-based lip-sync. To this end, we propose a novel lip-sync pipeline that, unlike previous methods, decouples the original chin contour from the lip-synced chin contour, and reduces flickering near the mouth.

3D人脸口型同步生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。