arXiv:2512.05121cs.GRcs.AI2025-12被引 7

用语音生成带个人情绪风格的3D人脸动画

PESTalk: Speech-Driven 3D Facial Animation with Personalized Emotional Styles

  • 双流音频特征提取+声纹建模,精准捕捉情绪
  • 在新构建的数据集上实现更逼真个性化的动画
  • 适合影视配音、虚拟人等需要情绪表达的场景

PESTalk 是一种直接从语音生成带个性化情绪风格的3D人脸动画的新方法。它克服了现有方法的局限,提出双流情绪提取器(DSEE),同时捕捉音频的时间与频域特征,实现细粒度情绪分析;并引入情绪风格建模模块(ESMM),基于声纹特征建模个体表达模式。为缓解数据稀缺问题,该方法构建了一个新的3D-EmoStyle数据集。评估结果表明,PESTalk 在生成逼真且个性化的面部动画方面优于当前最先进的方法。

原文摘要 · Abstract (English)

PESTalk is a novel method for generating 3D facial animations with personalized emotional styles directly from speech. It overcomes key limitations of existing approaches by introducing a Dual-Stream Emotion Extractor (DSEE) that captures both time and frequency-domain audio features for fine-grained emotion analysis, and an Emotional Style Modeling Module (ESMM) that models individual expression patterns based on voiceprint characteristics. To address data scarcity, the method leverages a newly constructed 3D-EmoStyle dataset. Evaluations demonstrate that PESTalk outperforms state-of-the-art methods in producing realistic and personalized facial animations.

3D动画语音驱动情绪建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。