构建40类细粒度情绪数据集,推动AI更准确理解人类复杂情感。
EmoNet-Face: An Expert-Annotated Benchmark for Synthetic Emotion Recognition
- 提出40类情绪分类体系,覆盖细微情感差异
- 创建三套可控生成数据集,含全脸表情与多元人口分布
- 提供专家标注基准,适合情感计算研究者使用
高效人机交互依赖于AI准确感知和理解人类情绪。现有视觉与多模态模型的评估基准严重受限,仅涵盖狭窄的情绪范围,忽略如苦涩、陶醉等细微状态,且难以区分相似情绪(如羞耻与尴尬)。现有数据集常使用未控图像,存在面部遮挡问题,缺乏人口多样性,易引入偏差。为解决这些关键缺陷,我们推出EmoNet Face——一个全面的基准套件。包含:(1) 基于基础研究构建的40类情绪分类体系,捕捉人类情感体验的细微差别;(2) 三个大规模生成数据集(EmoNet HQ、Binary 和 Big),具明确完整面部表情及跨种族、年龄、性别的均衡分布;(3) 多专家严格标注,用于训练与高保真评估;(4) 构建EmpathicInsight-Face模型,在该基准上达到人类专家水平。公开发布的EmoNet Face套件——包括分类体系、数据集与模型——为开发和评估具备深层情感理解能力的AI系统提供了坚实基础。
原文摘要 · Abstract (English)
Effective human-AI interaction relies on AI's ability to accurately perceive and interpret human emotions. Current benchmarks for vision and vision-language models are severely limited, offering a narrow emotional spectrum that overlooks nuanced states (e.g., bitterness, intoxication) and fails to distinguish subtle differences between related feelings (e.g., shame vs. embarrassment). Existing datasets also often use uncontrolled imagery with occluded faces and lack demographic diversity, risking significant bias. To address these critical gaps, we introduce EmoNet Face, a comprehensive benchmark suite. EmoNet Face features: (1) A novel 40-category emotion taxonomy, meticulously derived from foundational research to capture finer details of human emotional experiences. (2) Three large-scale, AI-generated datasets (EmoNet HQ, Binary, and Big) with explicit, full-face expressions and controlled demographic balance across ethnicity, age, and gender. (3) Rigorous, multi-expert annotations for training and high-fidelity evaluation. (4) We built EmpathicInsight-Face, a model achieving human-expert-level performance on our benchmark. The publicly released EmoNet Face suite - taxonomy, datasets, and model - provides a robust foundation for developing and evaluating AI systems with a deeper understanding of human emotions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。