arXiv:2502.20323cs.CV2025-02SIGGRAPH被引 27

用自回归模型实现语音驱动的实时3D头像动画,效果更自然。

ARTalk: Speech-Driven 3D Head Animation via Autoregressive Model

  • 基于自回归架构,从语音映射到多尺度动作码本生成动画。
  • 实时生成高度同步的口型、头部姿态和眨眼动作,准确率领先。
  • 可适配未见说话风格,适合个性化虚拟人创作。

语音驱动的3D面部动画旨在从任意音频片段生成逼真的口型和面部表情。尽管现有基于扩散模型的方法能生成自然动作,但其生成速度慢限制了应用潜力。本文提出一种新型自回归模型,通过学习语音到多尺度运动码本的映射,实现实时生成高度同步的口型、真实头部姿态和眼睑眨动。此外,该模型可适应未见的说话风格,使创建超越训练身份的独特个人风格3D说话形象成为可能。大量评估和用户研究显示,该方法在口型同步精度和感知质量上均优于现有技术。

原文摘要 · Abstract (English)

Speech-driven 3D facial animation aims to generate realistic lip movements and facial expressions for 3D head models from arbitrary audio clips. Although existing diffusion-based methods are capable of producing natural motions, their slow generation speed limits their application potential. In this paper, we introduce a novel autoregressive model that achieves real-time generation of highly synchronized lip movements and realistic head poses and eye blinks by learning a mapping from speech to a multi-scale motion codebook. Furthermore, our model can adapt to unseen speaking styles, enabling the creation of 3D talking avatars with unique personal styles beyond the identities seen during training. Extensive evaluations and user studies demonstrate that our method outperforms existing approaches in lip synchronization accuracy and perceived quality.

3D动画语音驱动自回归模型虚拟人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。