arXiv:2507.06071cs.CVcs.MM2025-07被引 11

让3D人脸动画随语音动态表达情绪,支持文本和图片控制。

MEDTalk: Multimodal Controlled 3D Facial Animation with Dynamic Emotions by Disentangled Embedding

  • 分离语音内容与情绪特征,实现唇动与表情独立控制。
  • 融合音频与文本,动态生成帧级情绪强度变化。
  • 支持文本描述和参考图引导,适合影视特效个性化生成。

基于音频驱动的情绪化3D人脸动画旨在生成同步的口型动作和生动的表情。然而,现有方法多依赖静态预设的情绪标签,限制了表达的多样性和自然性。为此,我们提出MEDTalk,一种细粒度且动态的情绪化说话头生成框架。该方法通过精心设计的交叉重建过程,从运动序列中解耦内容与情绪嵌入空间,实现对口型动作和面部表情的独立控制。除了传统的音频驱动口型同步外,还融合音频与语音文本,预测帧级情绪强度变化,并动态调整静态情绪特征以生成真实的情绪表达。此外,为增强可控性与个性化,引入多模态输入——包括文本描述和参考表情图像——以引导生成用户指定的面部表情。以MetaHuman为优先目标,生成结果可直接接入工业生产流程。代码已公开:https://github.com/SJTU-Lucy/MEDTalk。

原文摘要 · Abstract (English)

Audio-driven emotional 3D facial animation aims to generate synchronized lip movements and vivid facial expressions. However, most existing approaches focus on static and predefined emotion labels, limiting their diversity and naturalness. To address these challenges, we propose MEDTalk, a novel framework for fine-grained and dynamic emotional talking head generation. Our approach first disentangles content and emotion embedding spaces from motion sequences using a carefully designed cross-reconstruction process, enabling independent control over lip movements and facial expressions. Beyond conventional audio-driven lip synchronization, we integrate audio and speech text, predicting frame-wise intensity variations and dynamically adjusting static emotion features to generate realistic emotional expressions. Furthermore, to enhance control and personalization, we incorporate multimodal inputs-including text descriptions and reference expression images-to guide the generation of user-specified facial expressions. With MetaHuman as the priority, our generated results can be conveniently integrated into the industrial production pipeline. The code is available at: https://github.com/SJTU-Lucy/MEDTalk.

3D人脸动画情绪生成多模态控制音频驱动

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。