arXiv:2605.07478cs.CV2026-05

用语言信息提升语音驱动面部动画的准确性

AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models

论文配图:AudioFace: Language-Assisted Speech-Driven Facial Animation with Multimodal Language Models
图 1 · 摘自论文原文
  • 引入语义和发音层级线索,指导口部动作生成
  • 在多个指标上优于现有方法,尤其在嘴部动作同步性上表现突出
  • 适合需要高精度口型同步的应用场景

语音驱动面部动画需精确匹配声学信号与面部运动,尤其是与发音相关的口部动作。然而,直接将语音音频映射到面部系数常忽略语音产生的语言与音位结构。本文提出AudioFace,一种基于语言辅助的语音驱动混合形状生成框架,将口部相关面部系数预测视为受语言与发音信息引导的结构化生成问题。不依赖单一声学特征,该方法利用多模态大语言模型的先验知识,引入文本与音素级提示,建立语音信号与可解释面部动作之间的桥梁。大量实验表明,AudioFace在多个评估指标上表现更优,验证了语言辅助与多模态先验引导在语音驱动面部动画中的有效性。

原文摘要 · Abstract (English)

Speech-driven facial animation requires accurate correspondence between acoustic signals and facial motion, especially for articulation-related mouth movements. However, directly mapping speech audio to facial coefficients often overlooks the linguistic and phonetic structure underlying speech production. In this paper, we propose AudioFace, a language-assisted framework for speech-driven blendshape generation that treats mouth-related facial coefficient prediction as a structured generation problem guided by linguistic and articulatory information. Instead of relying solely on acoustic features, our method leverages the prior knowledge of multimodal large language models and introduces transcript- and phoneme-level cues to bridge speech signals with interpretable facial actions. Extensive experiments show that AudioFace achieves superior performance across multiple evaluation metrics, validating the effectiveness of language-assisted and multimodal-prior-guided speech-driven facial animation.

语音驱动面部动画多模态语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。