评测三大模型零样本情绪分类能力,发现均难识别爱、困惑等复杂情绪。
Quantifying the Affective Gap: A Zero-Shot Evaluation of LLMs on Fine-Grained Emotion Taxonomies

- 统一提示零样本测试,用13类细粒度情绪数据集评估模型表现。
- 谷歌Gemini准确率最高(39.9%),但所有模型对爱、羞耻等情绪识别差。
- 三模型差异不显著,显示当前大模型在情绪理解上存在共性瓶颈。
情感识别是情感计算中的基础挑战,对人机交互、心理健康支持和对话AI具有重要意义。本文对三种主流商业大模型——Claude(claude-sonnet-4-6)、ChatGPT(GPT-5.4)和Gemini(gemini-2.5-flash)——进行了统一的零样本评估。模型通过各自生产API于2026年4月进行测试,任务为细粒度13类情绪分类。使用来自boltuix/emotions数据集的1,000句分层采样样本(共131,306句,涵盖13类情绪),采用无示例的统一提示。Gemini准确率最高(39.9%),宏平均F1为0.363;GPT-5.4为38.8%(F1=0.291);Claude为38.0%(F1=0.159)。所有模型在讽刺与欲望类情绪表现良好,但在爱、困惑与羞耻上持续表现不佳。麦克内马尔检验显示各模型间无显著差异(p > 0.10),表明零样本情绪识别已达到共同上限。Claude的宏F1显著偏低,暴露其对类别不平衡的预测偏差。结果揭示前沿AI系统在零样本细粒度情绪识别上的局限性。
原文摘要 · Abstract (English)
Emotion recognition in natural language is a foundational challenge in affective computing, with critical implications for human-computer interaction, mental health support, and conversational AI. This paper presents a rigorous, unified zero-shot evaluation of three leading commercial large language models: Claude (claude-sonnet-4-6), ChatGPT (GPT-5.4), and Gemini (gemini-2.5-flash). The models were queried through their respective production APIs as of April 2026 on a fine-grained 13-class emotion classification task. Using a stratified 1,000-sentence sample from the boltuix/emotions dataset, which comprises 131,306 sentences across 13 categories, a single uniform prompt with no exemplars was applied identically across all models. Gemini achieves the highest accuracy (39.9%) and macro-F1 score (0.363), followed by GPT-5.4 (38.8%, macro-F1 = 0.291) and Claude (38.0%, macro-F1 = 0.159). All models excel on sarcasm and desire while consistently failing on love, confusion, and shame. McNemar tests reveal no statistically significant pairwise differences (p > 0.10), suggesting convergence at a shared zero-shot ceiling. Claude's markedly lower macro-F1 score exposes a class-imbalance prediction bias. These findings highlight the current limitations of frontier AI systems in zero-shot fine-grained emotion classification.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。