arXiv:2506.00900cs.CL2025-06ACL被引 13

构建双维度社交智能评估基准,揭示大模型在社交行为上的优劣与人类差异。

SocialEval: Evaluating Social Intelligence of Large Language Models

  • 基于叙事脚本构建双评价体系,兼顾结果与人际互动过程。
  • 大模型在社交目标达成和人际能力上均弱于人类,且偏好积极但低效行为。
  • 模型内部表征具功能分区,类似人脑,体现特定社交能力学习。

大型语言模型在模拟人类行为方面展现出可观的社交智能(SI),亟需对其社交智能及与人类的差异进行评估。社交智能使人能够在社会互动中智慧地行动以实现社交目标。为此,我们提出SocialEval——一个基于脚本的双语社交智能评估基准,通过人工设计叙事脚本,融合结果导向的目标达成评估与过程导向的人际能力评估。每个脚本构成一个世界树结构,包含由人际能力驱动的情节线,全面呈现模型如何应对社交情境。实验表明,大模型在两项评估中均落后于人类,表现出亲社会倾向,即使这种倾向导致目标失败。对模型表征空间和神经激活的分析显示,其已发展出类人脑的能力特异性功能分区。

原文摘要 · Abstract (English)

LLMs exhibit promising Social Intelligence (SI) in modeling human behavior, raising the need to evaluate LLMs' SI and their discrepancy with humans. SI equips humans with interpersonal abilities to behave wisely in navigating social interactions to achieve social goals. This presents an operational evaluation paradigm: outcome-oriented goal achievement evaluation and process-oriented interpersonal ability evaluation, which existing work fails to address. To this end, we propose SocialEval, a script-based bilingual SI benchmark, integrating outcome- and process-oriented evaluation by manually crafting narrative scripts. Each script is structured as a world tree that contains plot lines driven by interpersonal ability, providing a comprehensive view of how LLMs navigate social interactions. Experiments show that LLMs fall behind humans on both SI evaluations, exhibit prosociality, and prefer more positive social behaviors, even if they lead to goal failure. Analysis of LLMs' formed representation space and neuronal activations reveals that LLMs have developed ability-specific functional partitions akin to the human brain.

社交智能评估基准大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。