arXiv:2409.19143cs.CV2024-09被引 2

让同一段语音生成多种自然表情,提升面部动画多样性。

Diverse Code Query Learning for Speech-Driven Facial Animation

  • 通过多样化损失引导模型探索表情潜空间,生成多组不同动作
  • 在小数据集上仍能覆盖多样且合理的面部运动模式
  • 支持分部位逐次生成,实现可控的面部动画合成

语音驱动的面部动画旨在根据给定语音信号生成唇形同步的3D说话人脸。以往方法多聚焦于真实感,却较少关注面部动作的潜在随机性。虽然生成模型可通过重复采样处理一对多映射,但在小规模数据集上确保多样且合理的动作覆盖仍具挑战。本文提出:对同一音频信号预测多个样本,并显式鼓励样本多样性。核心思路是利用多样性促进损失引导模型探索表达丰富的面部潜空间,从而识别出理想的多样化潜码。基于向量量化变分自编码机制所学得的丰富面部先验,模型时序查询多个随机潜码,可灵活解码为多样且忠实于语音的面部动作。为进一步实现局部控制,模型采用顺序方式预测各面部区域并组合成完整表情动作。该框架统一实现了多样性和可控性。实验表明,本方法在定量与定性评估中均达到当前最优,尤其在样本多样性方面表现突出。

原文摘要 · Abstract (English)

Speech-driven facial animation aims to synthesize lip-synchronized 3D talking faces following the given speech signal. Prior methods to this task mostly focus on pursuing realism with deterministic systems, yet characterizing the potentially stochastic nature of facial motions has been to date rarely studied. While generative modeling approaches can easily handle the one-to-many mapping by repeatedly drawing samples, ensuring a diverse mode coverage of plausible facial motions on small-scale datasets remains challenging and less explored. In this paper, we propose predicting multiple samples conditioned on the same audio signal and then explicitly encouraging sample diversity to address diverse facial animation synthesis. Our core insight is to guide our model to explore the expressive facial latent space with a diversity-promoting loss such that the desired latent codes for diversification can be ideally identified. To this end, building upon the rich facial prior learned with vector-quantized variational auto-encoding mechanism, our model temporally queries multiple stochastic codes which can be flexibly decoded into a diverse yet plausible set of speech-faithful facial motions. To further allow for control over different facial parts during generation, the proposed model is designed to predict different facial portions of interest in a sequential manner, and compose them to eventually form full-face motions. Our paradigm realizes both diverse and controllable facial animation synthesis in a unified formulation. We experimentally demonstrate that our method yields state-of-the-art performance both quantitatively and qualitatively, especially regarding sample diversity.

面部动画语音驱动生成模型多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。