arXiv:2508.14214cs.AI2025-08被引 1

大模型对情绪刺激的判断与人类高度一致,尤其在快乐情绪上最接近。

Large Language Models are Highly Aligned with Human Ratings of Emotional Stimuli

  • 用人类标注的情绪数据测试大模型评分,对比多模态输入
  • GPT-4o与人类相关性超0.9,快乐情绪最匹配,唤醒度差异较大
  • 模型判断更统一,适合需要稳定情绪判断的场景

情感深刻影响人类行为与认知,尤其在日常和高压情境中。将大语言模型(LLMs)融入生活时,需了解其如何评估情绪化刺激。为此,我们采集了多个主流大模型对已由人类标注情绪内容的词语与图像数据集的评分。结果显示,在相同任务下,GPT-4o在不同模态、刺激类型和多数评分量表上与人类参与者高度一致(许多情况相关系数 r ≥ 0.9)。但唤醒度评分一致性较低,而快乐情绪最为吻合。总体而言,大模型在五类情绪框架(快乐、愤怒、悲伤、恐惧、厌恶)中的表现优于二维维度(唤醒度与效价)结构。此外,大模型评分比人类更趋同。这些结果揭示了大模型理解情绪刺激的方式,并凸显生物智能与人工智能在关键行为领域的一致性与差异。

原文摘要 · Abstract (English)

Emotions exert an immense influence over human behavior and cognition in both commonplace and high-stress tasks. Discussions of whether or how to integrate large language models (LLMs) into everyday life (e.g., acting as proxies for, or interacting with, human agents), should be informed by an understanding of how these tools evaluate emotionally loaded stimuli or situations. A model's alignment with human behavior in these cases can inform the effectiveness of LLMs for certain roles or interactions. To help build this understanding, we elicited ratings from multiple popular LLMs for datasets of words and images that were previously rated for their emotional content by humans. We found that when performing the same rating tasks, GPT-4o responded very similarly to human participants across modalities, stimuli and most rating scales (r = 0.9 or higher in many cases). However, arousal ratings were less well aligned between human and LLM raters, while happiness ratings were most highly aligned. Overall LLMs aligned better within a five-category (happiness, anger, sadness, fear, disgust) emotion framework than within a two-dimensional (arousal and valence) organization. Finally, LLM ratings were substantially more homogenous than human ratings. Together these results begin to describe how LLM agents interpret emotional stimuli and highlight similarities and differences among biological and artificial intelligence in key behavioral domains.

情绪识别大模型对齐GPT-4o人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。