arXiv:2603.15512cs.CV2026-03

无需模板即可生成带情绪的3D说话人脸,支持任意网格拓扑。

FreeTalk: Emotional Topology-Free 3D Talking Heads

  • 分两阶段:先从语音预测稀疏特征点运动,再映射到任意网格。
  • 在未见身份和网格结构下仍保持高鲁棒性,性能接近专用模型。
  • 适用于原始扫描数据,无需配准或对应监督,适合真实场景应用。

语音驱动的3D面部动画发展迅速,但多数方法依赖注册模板网格,难以应用于任意拓扑的原始3D扫描。同时,超越口型变化的情感动态建模仍具挑战,且常受限于模板参数化。为此,本文提出FreeTalk,一种两阶段情感条件化的3D说话头动画框架,可泛化至无注册、顶点数量与连接关系任意的面部网格。首先,音频到稀疏(ATS)模块从语音中预测时序连贯的3D特征点位移序列,受情感类别与强度控制,该稀疏表示捕捉了发音与情感运动,且与网格拓扑无关。其次,稀疏到网格(STM)模块结合表面内在特征与特征点-顶点条件,将预测位移转移到目标网格,生成稠密顶点变形,测试时无需模板拟合或对应监督。大量实验表明,FreeTalk在域内训练时性能媲美专用基线,而在未见身份与网格拓扑下显著提升鲁棒性。代码与预训练模型将公开发布。

原文摘要 · Abstract (English)

Speech-driven 3D facial animation has advanced rapidly, yet most approaches remain tied to registered template meshes, preventing effective deployment on raw 3D scans with arbitrary topology. At the same time, modeling controllable emotional dynamics beyond lip articulation remains challenging, and is often tied to template-based parameterizations. We address these challenges by proposing FreeTalk, a two-stage framework for emotion-conditioned 3D talking-head animation that generalizes to unregistered face meshes with arbitrary vertex count and connectivity. First, Audio-To-Sparse (ATS) predicts a temporally coherent sequence of 3D landmark displacements from speech audio, conditioned on an emotion category and intensity. This sparse representation captures both articulatory and affective motion while remaining independent of mesh topology. Second, Sparse-To-Mesh (STM) transfers the predicted landmark motion to a target mesh by combining intrinsic surface features with landmark-to-vertex conditioning, producing dense per-vertex deformations without template fitting or correspondence supervision at test time. Extensive experiments show that FreeTalk matches specialized baselines when trained in-domain, while providing substantially improved robustness to unseen identities and mesh topologies. Code and pre-trained models will be made publicly available.

3D说话头情感动画非模板化拓扑自由

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。