构建首个表情符号提示的动态人脸表情数据集,评估模型对多样感知的理解能力
Chehre: An Emoji-Prompted Video Dataset for Perceptually Diverse Facial Expression Recognition

- 用表情符号引导拍摄40种动态面部表情,保护隐私并保留动作特征
- 2111段视频来自203人,902人标注,最佳模型在主要表情识别上仅达32.5%准确率
- 提出两种新任务:主导表情与分布式表情识别,适合研究人类感知多样性
面部表情是人际互动中的非语言社交信号,但现有表情识别数据集多聚焦静态图像、基础情绪类别或单一确定性标注。我们提出Chehre,一个由表情符号提示的视频数据集,用于分析跨广泛表情的动态面部表情及其个体间感知差异。参与者根据40个表情符号录制面部动作,随后将真实面部运动迁移至合成人脸以保障隐私。另一组标注者对匿名化视频进行表情符号与标签标注,最终收集到2111段高质量视频,涵盖203名表演者,经902名标注者验证。我们定义两个基准任务:主导表情识别(测试模型是否能恢复人类评分最高的标签)和分布式表情识别(测试模型是否捕捉人类响应的多样性)。使用随机采样与角色提示生成每视频多个预测结果,对近期视觉-语言模型进行评测。结果显示两项任务均具挑战性:在主导表情识别中,表现最佳模型仅有32.5%的Top-1准确率;在分布式识别中,模型的Spread Ratio远低于人类参考水平。Chehre为评估多样化、动态化、分布式的面部表情识别提供了基准。
原文摘要 · Abstract (English)
Facial expressions are nonverbal social signals used in human interaction, but facial expression recognition datasets often focus on static images, basic emotion categories, or single deterministic annotations. We introduce Chehre, an emoji-prompted video dataset for analyzing dynamic facial expressions across a wide range of expressions for exploring inter-individual perceptual diversity. In Chehre, participants were prompted to express and record 40 facial emojis. Later, their facial motions were transferred onto synthetic faces to preserve privacy. A separate group of annotators analyzed the anonymized videos using emoji and label annotations, resulting in 2,111 high quality videos collected from 203 performers and validated by 902 annotators. We define two benchmark tasks: dominant expression recognition, which tests whether models recover the top human-rated labels, and distributional expression recognition, which tests whether models capture the diversity of human responses. We benchmark recent vision-language models using random sampling and persona prompting to generate multiple predictions per video. Results show that both tasks are challenging: among the models evaluated, the best-performing model achieves only 32.5% Top-1 accuracy on dominant expression recognition and a Spread Ratio well below the human reference on distributional recognition. Chehre provides a benchmark for evaluating diverse, dynamic, and distributional facial expression recognition
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。