构建首个图文表情生成数据集,提升情绪驱动的3D人脸动画质量
Emo3D: Metric and Benchmarking Dataset for 3D Facial Expression Generation from Emotion Description
- 用大模型生成多样情绪描述,配以3D混合形状和图像
- 新指标比MSE更准确评估表情与文本的情感对齐度
- 适合虚拟人、动画设计及情感交互系统研究者
现有3D人脸情绪建模受限于情绪类别少和数据不足。本文提出「Emo3D」,一个涵盖广泛人类情绪的「文本-图像-表情」数据集,每种情绪均配有图像与3D混合形状。借助大语言模型(LLMs)生成多样化文本描述,实现对丰富情绪表达的捕捉。基于该数据集,我们全面评估了语言模型微调及视觉语言模型(如CLIP)在3D面部表情合成中的表现。同时引入新评价指标——Emo3D,其在衡量3D表情与人类情绪相关联的视觉-文本对齐度和语义丰富性方面优于均方误差(MSE)。Emo3D可广泛应用于动画设计、虚拟现实及情感人机交互领域。
原文摘要 · Abstract (English)
Existing 3D facial emotion modeling have been constrained by limited emotion classes and insufficient datasets. This paper introduces "Emo3D", an extensive "Text-Image-Expression dataset" spanning a wide spectrum of human emotions, each paired with images and 3D blendshapes. Leveraging Large Language Models (LLMs), we generate a diverse array of textual descriptions, facilitating the capture of a broad spectrum of emotional expressions. Using this unique dataset, we conduct a comprehensive evaluation of language-based models' fine-tuning and vision-language models like Contranstive Language Image Pretraining (CLIP) for 3D facial expression synthesis. We also introduce a new evaluation metric for this task to more directly measure the conveyed emotion. Our new evaluation metric, Emo3D, demonstrates its superiority over Mean Squared Error (MSE) metrics in assessing visual-text alignment and semantic richness in 3D facial expressions associated with human emotions. "Emo3D" has great applications in animation design, virtual reality, and emotional human-computer interaction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。