arXiv:2505.01932cs.GRcs.CV2025-05被引 9

用最优传输优化3D说话头动画,让口型更准动作更自然。

OT-Talk: Animating 3D Talking Head with Optimal Transportation

  • 引入最优传输思想,用切比雪夫图卷积提取网格几何特征。
  • 在两个公开数据集上,重建精度和时序对齐均优于现有方法。
  • 适合做虚拟主播、游戏角色等需要逼真口型同步的场景。

基于音频驱动3D头部网格动画在AR/VR、游戏和娱乐中应用广泛,但语音信号与面部动态之间的模态差异导致口型不同步和动作不自然。为解决此问题,本文提出OT-Talk,首个利用最优传输优化学习模型的说话头动画方法。在现有框架基础上,采用预训练Hubert模型提取音频特征,使用Transformer处理时序序列。不同于仅关注顶点坐标或位移的方法,我们引入切比雪夫图卷积从三角网格中提取几何特征。为衡量网格差异,我们超越传统重建误差和相邻帧速度差,将网格表示为概率测度并近似其表面,从而利用切片沃尔什距离建模网格变化。该方法使面部运动更平滑准确,生成连贯自然的动画。在两个公开音频-网格数据集上的实验表明,本方法在重建精度和时间对齐方面均优于当前最优技术。此外,20名志愿者参与的用户感知研究进一步验证了其有效性。

原文摘要 · Abstract (English)

Animating 3D head meshes using audio inputs has significant applications in AR/VR, gaming, and entertainment through 3D avatars. However, bridging the modality gap between speech signals and facial dynamics remains a challenge, often resulting in incorrect lip syncing and unnatural facial movements. To address this, we propose OT-Talk, the first approach to leverage optimal transportation to optimize the learning model in talking head animation. Building on existing learning frameworks, we utilize a pre-trained Hubert model to extract audio features and a transformer model to process temporal sequences. Unlike previous methods that focus solely on vertex coordinates or displacements, we introduce Chebyshev Graph Convolution to extract geometric features from triangulated meshes. To measure mesh dissimilarities, we go beyond traditional mesh reconstruction errors and velocity differences between adjacent frames. Instead, we represent meshes as probability measures and approximate their surfaces. This allows us to leverage the sliced Wasserstein distance for modeling mesh variations. This approach facilitates the learning of smooth and accurate facial motions, resulting in coherent and natural facial animations. Our experiments on two public audio-mesh datasets demonstrate that our method outperforms state-of-the-art techniques both quantitatively and qualitatively in terms of mesh reconstruction accuracy and temporal alignment. In addition, we conducted a user perception study with 20 volunteers to further assess the effectiveness of our approach.

3D动画说话头最优传输口型同步

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。