从第一人称视角系统评估大模型社会智能,发现顶尖模型仍不及人类。
EgoSocialArena: Benchmarking the Social Intelligence of Large Language Models from a First-person Perspective
- 构建第一人称视角评测框架,覆盖认知、情境、行为三方面智能。
- 8个主流大模型测试中,连O1-preview也未达人类水平。
- 聚焦人机交互场景,填补行为智能评估空白,适合智能体研发者参考。
社会智能建立在认知智能、情境智能和行为智能三大支柱之上。随着大语言模型(LLMs)日益融入社会生活,理解、评估和发展其社会智能变得愈发重要。现有研究多聚焦单一维度,且通常将模型置于第三人称观察视角(如心智理论测试),难以匹配真实智能体使用场景。相比之下,以第一人称为中心的评估更贴近实际应用。此外,行为智能缺乏全面评估,尤其缺少关键人机交互场景的考量。为此,我们提出EgoSocialArena——一个基于社会智能三支柱的新型评测框架,旨在系统性地从第一人称视角评估大模型的社会智能。通过该框架,我们对八种主流基础模型进行了全面评估,结果显示,即便是最先进的模型如O1-preview,仍显著落后于人类表现。
原文摘要 · Abstract (English)
Social intelligence is built upon three foundational pillars: cognitive intelligence, situational intelligence, and behavioral intelligence. As large language models (LLMs) become increasingly integrated into our social lives, understanding, evaluating, and developing their social intelligence are becoming increasingly important. While multiple existing works have investigated the social intelligence of LLMs, (1) most focus on a specific aspect, and the social intelligence of LLMs has yet to be systematically organized and studied; (2) position LLMs as passive observers from a third-person perspective, such as in Theory of Mind (ToM) tests. Compared to the third-person perspective, ego-centric first-person perspective evaluation can align well with actual LLM-based Agent use scenarios. (3) a lack of comprehensive evaluation of behavioral intelligence, with specific emphasis on incorporating critical human-machine interaction scenarios. In light of this, we present EgoSocialArena, a novel framework grounded in the three pillars of social intelligence: cognitive, situational, and behavioral intelligence, aimed to systematically evaluate the social intelligence of LLMs from a first-person perspective. With EgoSocialArena, we conduct a comprehensive evaluation of eight prominent foundation models, even the most advanced LLMs like O1-preview lag behind human performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。