arXiv:2507.20562cs.CVcs.AI2025-07ICCV被引 9

仅用音频生成个性化3D人脸动画,无需额外信息。

MemoryTalker: Personalized Speech-Driven 3D Facial Animation via Audio-Guided Stylization

  • 通过音频驱动的风格化记忆库,实现无先验的个性化表情合成。
  • 在多种数据集上优于现有方法,生成更自然、贴合说话风格的面部动作。
  • 适合需要真实互动式虚拟角色的应用场景。

语音驱动的3D人脸动画旨在从给定音频中合成逼真的人脸运动序列,匹配说话者的表达风格。然而,以往方法通常依赖说话者类别标签或推理时的额外3D人脸网格,难以体现真实说话风格,限制了实际应用。为此,我们提出MemoryTalker,仅通过音频输入即可实现高保真且个性化的3D人脸运动合成,最大化应用场景的实用性。框架包含两个训练阶段:第一阶段为存储和检索通用动作(即“记忆”),第二阶段利用音频驱动的说话风格特征对动作记忆进行风格化,实现个性化动画生成(即“动画化”)。该阶段学习特定音频应强调哪些面部动作类型。实验表明,我们的模型在定量评估、定性对比及用户研究中均表现优异,显著提升个性化人脸动画效果,优于当前最优方法。

原文摘要 · Abstract (English)

Speech-driven 3D facial animation aims to synthesize realistic facial motion sequences from given audio, matching the speaker's speaking style. However, previous works often require priors such as class labels of a speaker or additional 3D facial meshes at inference, which makes them fail to reflect the speaking style and limits their practical use. To address these issues, we propose MemoryTalker which enables realistic and accurate 3D facial motion synthesis by reflecting speaking style only with audio input to maximize usability in applications. Our framework consists of two training stages: 1-stage is storing and retrieving general motion (i.e., Memorizing), and 2-stage is to perform the personalized facial motion synthesis (i.e., Animating) with the motion memory stylized by the audio-driven speaking style feature. In this second stage, our model learns about which facial motion types should be emphasized for a particular piece of audio. As a result, our MemoryTalker can generate a reliable personalized facial animation without additional prior information. With quantitative and qualitative evaluations, as well as user study, we show the effectiveness of our model and its performance enhancement for personalized facial animation over state-of-the-art methods.

3D动画语音驱动个性化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。