测试大模型情感认知能力,发现其表现媲美甚至超越人类。
Human-like Affective Cognition in Foundation Models
- 基于心理学理论构建1280个情感场景,评估模型推理能力。
- GPT-4等模型在多数情境下与人类判断一致,部分超人类平均水平。
- 思维链提示显著提升表现,表明模型具备类人情感理解力。
理解情绪是人类互动与体验的核心。人类能轻松从情境或表情推断情绪,或由情绪反推情境,完成多种情感认知。当前人工智能在这些推断上表现如何?我们提出一个评估框架,基于心理学理论生成1280个多样化情境,涵盖评估、情绪、表情与结果之间的关系。在精心设计的条件下,评估GPT-4、Claude-3、Gemini-1.5-Pro等基础模型以及567名人类参与者的表现。结果显示,基础模型普遍与人类直觉一致,匹配或超过人类间的共识水平。在某些条件下,模型表现“超人”——比平均人类更准确预测主流判断。所有模型均受益于思维链推理。这表明基础模型已习得类似人类的情绪理解及其对信念与行为的影响。
原文摘要 · Abstract (English)
Understanding emotions is fundamental to human interaction and experience. Humans easily infer emotions from situations or facial expressions, situations from emotions, and do a variety of other affective cognition. How adept is modern AI at these inferences? We introduce an evaluation framework for testing affective cognition in foundation models. Starting from psychological theory, we generate 1,280 diverse scenarios exploring relationships between appraisals, emotions, expressions, and outcomes. We evaluate the abilities of foundation models (GPT-4, Claude-3, Gemini-1.5-Pro) and humans (N = 567) across carefully selected conditions. Our results show foundation models tend to agree with human intuitions, matching or exceeding interparticipant agreement. In some conditions, models are ``superhuman'' -- they better predict modal human judgements than the average human. All models benefit from chain-of-thought reasoning. This suggests foundation models have acquired a human-like understanding of emotions and their influence on beliefs and behavior.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。