从文本生成多样且自然的情感表情,提升数字人表现力。
When Words Smile: Generating Diverse Emotional Facial Expressions from Text
- 在连续潜在空间中学习情感动态,实现表情生成
- 在15,000对文本-3D表情数据上超越基线模型
- 适合对话系统、游戏等需要真实情感表达的场景
让数字人具备丰富情感表达在对话系统、游戏等交互场景中具有重要意义。尽管近期说话头合成技术在唇形同步方面取得显著进展,但往往忽视了面部表情的丰富性与动态性。为填补这一关键空白,我们提出一个端到端的文本到表情生成模型,明确聚焦于情感动态建模。该模型在连续潜在空间中学习表情变化,生成多样、流畅且情感一致的表情。为此,我们构建了EmoAva数据集,包含15,000组高质量的文本-3D表情配对。在现有数据集及EmoAva上的大量实验表明,我们的方法在多个评估指标上显著优于基线模型,标志着该领域的重要进展。
原文摘要 · Abstract (English)
Enabling digital humans to express rich emotions has significant applications in dialogue systems, gaming, and other interactive scenarios. While recent advances in talking head synthesis have achieved impressive results in lip synchronization, they tend to overlook the rich and dynamic nature of facial expressions. To fill this critical gap, we introduce an end-to-end text-to-expression model that explicitly focuses on emotional dynamics. Our model learns expressive facial variations in a continuous latent space and generates expressions that are diverse, fluid, and emotionally coherent. To support this task, we introduce EmoAva, a large-scale and high-quality dataset containing 15,000 text-3D expression pairs. Extensive experiments on both existing datasets and EmoAva demonstrate that our method significantly outperforms baselines across multiple evaluation metrics, marking a significant advancement in the field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。