构建可商用设备采集的4D人脸表情数据集,支持细粒度语言控制
Express4D: Expressive, Friendly, and Extensible 4D Facial Motion Generation Benchmark
- 用普通设备+LLM生成指令采集带语义标注的表情序列
- 提供丰富表达细节,支持文本到表情的多对多映射学习
- 适合动画、虚拟形象等需要精细表情控制的场景
从自然语言生成动态面部表情是计算机图形学中的关键任务,广泛应用于动画、虚拟化身和人机交互。然而现有生成模型依赖语音驱动或粗粒度情绪标签的数据集,缺乏细腻的表达描述,且采集需昂贵设备。为此,我们提出一个新数据集,包含细腻表演的人脸动作序列及语义标注,可通过消费级设备与大模型生成的自然语言指令轻松采集,输出为ARKit blendshape格式,具备可驱动性、丰富表达与标签信息。我们基于该数据集训练两个基线模型,并评估其性能以供未来研究参考。实验表明,模型能学习有意义的文到表情运动生成,捕捉两模态间的多对多映射关系。数据集、代码及示例视频已公开于:https://jaron1990.github.io/Express4D/
原文摘要 · Abstract (English)
Dynamic facial expression generation from natural language is a crucial task in Computer Graphics, with applications in Animation, Virtual Avatars, and Human-Computer Interaction. However, current generative models suffer from datasets that are either speech-driven or limited to coarse emotion labels, lacking the nuanced, expressive descriptions needed for fine-grained control, and were captured using elaborate and expensive equipment. We hence present a new dataset of facial motion sequences featuring nuanced performances and semantic annotation. The data is easily collected using commodity equipment and LLM-generated natural language instructions, in the popular ARKit blendshape format. This provides riggable motion, rich with expressive performances and labels. We accordingly train two baseline models, and evaluate their performance for future benchmarking. Using our Express4D dataset, the trained models can learn meaningful text-to-expression motion generation and capture the many-to-many mapping of the two modalities. The dataset, code, and video examples are available on our webpage: https://jaron1990.github.io/Express4D/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。