arXiv:2604.02941cs.CV2026-04

用多分辨率与多模态融合提升语音驱动3D人脸动画的同步精度。

MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion

论文配图:MMTalker: Multiresolution 3D Talking Head Synthesis with Multimodal Feature Fusion
图 1 · 摘自论文原文
  • 通过网格参数化与可学习采样实现面部细节连续表征。
  • 在唇动与眼动同步性上显著优于当前最佳方法。
  • 适合需要高保真语音驱动动画的影视与虚拟人应用。

语音驱动的三维面部动画合成旨在建立从一维语音信号到时变三维面部运动信号的映射。现有方法仍面临唇同步精度不足和表情不自然的问题,主要源于跨模态映射的高度病态性。本文提出一种新型3D音频驱动面部动画合成方法MMTalker,通过多分辨率表示与多模态特征融合,精准重建三维面部运动的丰富细节。首先,采用网格参数化与非均匀可微采样实现面部的连续表征,其中网格参数化建立UV平面与3D面部网格的对应关系,作为连续学习的真值;可微非均匀采样通过在每个三角面片设置可学习采样概率,实现精确的面部细节获取。其次,利用残差图卷积网络与双交叉注意力机制,从多输入模态中提取判别性面部运动特征,充分融合语音的层次特征与面部网格的显式时空几何特征。最后,轻量级回归网络联合处理规范化的UV空间采样点与编码后的面部运动特征,预测合成说话人脸的顶点级几何位移。大量实验表明,相比当前最优方法,该方法在唇部与眼部动作同步性方面取得显著提升。

原文摘要 · Abstract (English)

Speech-driven three-dimensional (3D) facial animation synthesis aims to build a mapping from one-dimensional (1D) speech signals to time-varying 3D facial motion signals. Current methods still face challenges in maintaining lip-sync accuracy and producing realistic facial expressions, primarily due to the highly ill-posed nature of this cross-modal mapping. In this paper, we introduce a novel 3D audio-driven facial animation synthesis method through multi-resolution representation and multi-modal feature fusion, called MMTalker which can accurately reconstruct the rich details of 3D facial motion. We first achieve the continuous representation of 3D face with details by mesh parameterization and non-uniform differentiable sampling. The mesh parameterization technique establishes the correspondence between UV plane and 3D facial mesh and is used to offer ground truth for the continuous learning. Differentiable non-uniform sampling enables precise facial detail acquisition by setting learnable sampling probability in each triangular face. Next, we employ residual graph convolutional network and dual cross-attention mechanism to extract discriminative facial motion feature from multiple input modalities. This proposed multimodal fusion strategy takes full use of the hierarchical features of speech and the explicit spatiotemporal geometric features of facial mesh. Finally, a lightweight regression network predicts the vertex-wise geometric displacements of the synthesized talking face by jointly processing the sampled points in the canonical UV space and the encoded facial motion features. Comprehensive experiments demonstrate that significant improvements are achieved over state-of-the-art methods, especially in the synchronization accuracy of lip and eye movements.

3D人脸动画语音驱动多模态融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。