arXiv:2505.01319cs.GRcs.LG2025-05International Conf…被引 7

用语音驱动面部动画,还能精准复制表演风格。

Model See Model Do: Speech-Driven Facial Animation with Style Control

  • 基于参考视频提取风格基元,引导扩散模型生成
  • 在保持口型同步的前提下,还原细微表演风格
  • 适合虚拟主播、游戏角色等需要个性化表达的场景

语音驱动的3D面部动画在虚拟形象、游戏和数字内容创作中至关重要。现有方法虽能实现准确的口型同步和基础表情生成,但难以捕捉并有效传递细腻的表演风格。我们提出一种新的基于样例的生成框架,通过参考风格视频条件化一个潜在扩散模型,生成高度表现力且时间连贯的面部动画。为解决风格参考的精确对齐问题,我们引入一种名为风格基元(style basis)的新条件机制,从参考视频中提取关键姿态,并以加性方式引导扩散过程,使生成结果贴合风格特征,同时不损害口型同步质量。大量定性、定量与感知评估表明,该方法在多种语音场景下均能忠实再现目标风格,且唇音同步性能更优。

原文摘要 · Abstract (English)

Speech-driven 3D facial animation plays a key role in applications such as virtual avatars, gaming, and digital content creation. While existing methods have made significant progress in achieving accurate lip synchronization and generating basic emotional expressions, they often struggle to capture and effectively transfer nuanced performance styles. We propose a novel example-based generation framework that conditions a latent diffusion model on a reference style clip to produce highly expressive and temporally coherent facial animations. To address the challenge of accurately adhering to the style reference, we introduce a novel conditioning mechanism called style basis, which extracts key poses from the reference and additively guides the diffusion generation process to fit the style without compromising lip synchronization quality. This approach enables the model to capture subtle stylistic cues while ensuring that the generated animations align closely with the input speech. Extensive qualitative, quantitative, and perceptual evaluations demonstrate the effectiveness of our method in faithfully reproducing the desired style while achieving superior lip synchronization across various speech scenarios.

语音驱动面部动画风格控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。