arXiv:2505.02847cs.CLcs.AI2025-05ACL被引 21

用模拟人类情感变化的虚拟角色评估大模型的高阶社交认知能力

Sentient Agent as a Judge: Evaluating Higher-Order Social Cognition in Large Language Models

  • 构建会随对话动态变化情绪与内心想法的虚拟角色,模拟真实人际互动
  • 实测显示模型情感轨迹与专业心理评估量表高度相关,验证评测可信度
  • 公开榜单揭示顶尖模型比旧模型情感理解力强4倍,传统评测未反映此差距

评估大语言模型(LLM)对人类而非仅文本的理解能力仍是一个开放挑战。为此,我们提出Sentient Agent as a Judge(SAGE),一种衡量LLM高阶社会认知的自动化评估框架。SAGE通过构建一个模拟人类情感变化和内在思维的虚拟角色,在多轮对话中实时推理其情绪变化、感受状态及应答策略,生成可解释的情感轨迹与内心独白。在100个支持性对话场景中的实验表明,最终的“感知情感评分”与Barrett-Lennard关系量表(BLRI)评分及逐句共情指标高度相关,验证了心理真实度。我们还构建了涵盖18个商用与开源模型的公开“感知领导者榜”,发现前沿系统(GPT-4o-Latest、Gemini2.5-Pro)与早期基线间存在高达4倍的差距,而这一差异在传统排行榜(如Arena)中未被体现。SAGE为追踪真正具备同理心与社交智能的语言代理发展提供了原则性强、可扩展且可解释的工具。

原文摘要 · Abstract (English)

Assessing how well a large language model (LLM) understands human, rather than merely text, remains an open challenge. To bridge the gap, we introduce Sentient Agent as a Judge (SAGE), an automated evaluation framework that measures an LLM's higher-order social cognition. SAGE instantiates a Sentient Agent that simulates human-like emotional changes and inner thoughts during interaction, providing a more realistic evaluation of the tested model in multi-turn conversations. At every turn, the agent reasons about (i) how its emotion changes, (ii) how it feels, and (iii) how it should reply, yielding a numerical emotion trajectory and interpretable inner thoughts. Experiments on 100 supportive-dialogue scenarios show that the final Sentient emotion score correlates strongly with Barrett-Lennard Relationship Inventory (BLRI) ratings and utterance-level empathy metrics, validating psychological fidelity. We also build a public Sentient Leaderboard covering 18 commercial and open-source models that uncovers substantial gaps (up to 4x) between frontier systems (GPT-4o-Latest, Gemini2.5-Pro) and earlier baselines, gaps not reflected in conventional leaderboards (e.g., Arena). SAGE thus provides a principled, scalable and interpretable tool for tracking progress toward genuinely empathetic and socially adept language agents.

社交认知情感评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。