arXiv:2409.10848cs.CVcs.AI2024-09被引 3

用动作控制生成更自然流畅的语音驱动3D人脸动画

3DFacePolicy: Audio-Driven 3D Facial Animation Based on Action Control

  • 将逐帧顶点生成改为基于动作序列的控制
  • 在VOCASET和BIWI数据集上显著优于现有方法
  • 适合需要高动态表情与流畅性的实时应用

语音驱动的3D人脸动画在研究和应用中已取得显著进展。尽管近期基线方法因逐帧生成顶点而难以实现自然连续的面部运动,我们提出3DFacePolicy,开创性地通过“动作”概念重新定义连续帧间顶点轨迹的变化。通过预测每个顶点的动作序列以编码帧间运动,将顶点生成重构为基于动作的控制范式。具体而言,我们采用扩散策略(diffusion policy)这一机器人控制机制,基于音频和顶点状态联合预测动作序列。在VOCASET和BIWI数据集上的大量实验表明,该方法显著优于当前最优方法,尤其擅长生成动态、富有表现力且自然平滑的面部动画。

原文摘要 · Abstract (English)

Audio-driven 3D facial animation has achieved significant progress in both research and applications. While recent baselines struggle to generate natural and continuous facial movements due to their frame-by-frame vertex generation approach, we propose 3DFacePolicy, a pioneer work that introduces a novel definition of vertex trajectory changes across consecutive frames through the concept of "action". By predicting action sequences for each vertex that encode frame-to-frame movements, we reformulate vertex generation approach into an action-based control paradigm. Specifically, we leverage a robotic control mechanism, diffusion policy, to predict action sequences conditioned on both audio and vertex states. Extensive experiments on VOCASET and BIWI datasets demonstrate that our approach significantly outperforms state-of-the-art methods and is particularly expert in dynamic, expressive and naturally smooth facial animations.

3D人脸动画语音驱动动作控制扩散模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。