arXiv:2501.01808cs.CV2025-01CVPR被引 20

用6种情绪专家模型,让人脸动画更真实表达复杂情感。

MoEE: Mixture of Emotion Experts for Audio-Driven Portrait Animation

  • 拆分6种基础情绪,实现单一与复合情感精准生成
  • 构建含150段视频的专用情感数据集,支持复杂情绪训练
  • 可仅凭音频控制情绪,适合影视/虚拟主播场景

说话头像生成已实现精确音频同步,但要生成逼真动态需捕捉丰富情感与细微表情。现有方法存在两大瓶颈:一是缺乏对单一基础情绪建模的框架,限制了复合情绪生成;二是缺少富含人类情感表达的完整数据集,制约模型潜力。为此,本文提出:1)混合情绪专家(MoEE)模型,将6种基础情绪解耦,实现单一与复合情绪的精准合成;2)专为情感驱动设计的DH-FaceEmoVid-150数据集,包含6种常见情绪及4类复合情绪,共150段视频,显著拓展模型训练能力。此外,引入情绪到潜在空间模块,融合音频、文本、标签等多模态输入,实现灵活的情绪控制,支持仅用音频操控情感。通过大量定量与定性评估,证明MoEE框架结合该数据集,在生成复杂情感与精细面部细节方面表现优异,树立新基准。相关数据集将公开发布。

原文摘要 · Abstract (English)

The generation of talking avatars has achieved significant advancements in precise audio synchronization. However, crafting lifelike talking head videos requires capturing a broad spectrum of emotions and subtle facial expressions. Current methods face fundamental challenges: a) the absence of frameworks for modeling single basic emotional expressions, which restricts the generation of complex emotions such as compound emotions; b) the lack of comprehensive datasets rich in human emotional expressions, which limits the potential of models. To address these challenges, we propose the following innovations: 1) the Mixture of Emotion Experts (MoEE) model, which decouples six fundamental emotions to enable the precise synthesis of both singular and compound emotional states; 2) the DH-FaceEmoVid-150 dataset, specifically curated to include six prevalent human emotional expressions as well as four types of compound emotions, thereby expanding the training potential of emotion-driven models. Furthermore, to enhance the flexibility of emotion control, we propose an emotion-to-latents module that leverages multimodal inputs, aligning diverse control signals-such as audio, text, and labels-to ensure more varied control inputs as well as the ability to control emotions using audio alone. Through extensive quantitative and qualitative evaluations, we demonstrate that the MoEE framework, in conjunction with the DH-FaceEmoVid-150 dataset, excels in generating complex emotional expressions and nuanced facial details, setting a new benchmark in the field. These datasets will be publicly released.

人脸动画情感生成多模态控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。