构建1225个互动社交场景,评测语言模型在复杂人际关系中的推理能力。
AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios
- 基于戏剧理论自底向上生成多样化社交场景
- 发现大模型在高阶成长需求场景中表现不足,私密信息推理仍需提升
- 适合研究智能体社会认知与人机交互的学者参考
大型语言模型(LLMs)正被广泛用于驱动自主智能体,在行为研究等领域模拟人类。然而,评估其在复杂社交互动中的能力仍具挑战。以往研究受限于场景多样性不足、复杂度低及单视角分析。为此,我们提出AgentSense:通过互动场景评测语言智能体的社会智力。基于戏剧理论,我们从大量剧本中构建了1,225个多样化的社交场景。通过多轮交互评估LLM驱动的智能体,重点关注目标达成与隐含推理。采用ERG理论分析目标,并开展全面实验。结果表明,大模型在复杂社交场景中难以实现高阶成长类目标,即便GPT-4o在私密信息推理方面仍有改进空间。代码与数据已开源。
原文摘要 · Abstract (English)
Large language models (LLMs) are increasingly leveraged to empower autonomous agents to simulate human beings in various fields of behavioral research. However, evaluating their capacity to navigate complex social interactions remains a challenge. Previous studies face limitations due to insufficient scenario diversity, complexity, and a single-perspective focus. To this end, we introduce AgentSense: Benchmarking Social Intelligence of Language Agents through Interactive Scenarios. Drawing on Dramaturgical Theory, AgentSense employs a bottom-up approach to create 1,225 diverse social scenarios constructed from extensive scripts. We evaluate LLM-driven agents through multi-turn interactions, emphasizing both goal completion and implicit reasoning. We analyze goals using ERG theory and conduct comprehensive experiments. Our findings highlight that LLMs struggle with goals in complex social scenarios, especially high-level growth needs, and even GPT-4o requires improvement in private information reasoning. Code and data are available at \url{https://github.com/ljcleo/agent_sense}.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。