构建长期社交互动评估基准,检验语言模型的社会智能持久性。
LIFELONG SOTOPIA: Evaluating Social Intelligence of Language Agents Over Lifelong Social Interactions
- 设计多轮角色扮演任务,模拟长期社交互动场景。
- 模型表现随交互轮次下降,历史理解能力不足导致目标完成率低。
- 适合研究社会智能、人机交互与长期记忆的学者参考。
人类在不同情境下与不同对象进行长期社交互动,需具备持续获取信息并有效应对多种社交环境的能力。当前研究对人工智能系统是否具备此类社会智能尚缺乏深入探讨。本文提出全新评估基准 LIFELONG-SOTOPIA,通过模拟多轮交互,全面评估语言代理在长期社交中的表现。每轮中,语言代理扮演特定角色,完成随机生成的社交任务。实验发现,所有测试模型在长期互动中目标达成率和角色可信度均持续下降;尽管采用先进记忆方法有所提升,最佳模型在需要显式理解互动历史的场景中仍显著低于人类表现。结果表明,LIFELONG-SOTOPIA 能有效评估语言代理在长期社交互动中的社会智能水平。
原文摘要 · Abstract (English)
Humans engage in lifelong social interactions through interacting with different people under different scenarios for different social goals. This requires social intelligence to gather information through a long time span and use it to navigate various social contexts effectively. Whether AI systems are also capable of this is understudied in the existing research. In this paper, we present a novel benchmark, LIFELONG-SOTOPIA, to perform a comprehensive evaluation of language agents by simulating multi-episode interactions. In each episode, the language agents role-play characters to achieve their respective social goals in randomly sampled social tasks. With LIFELONG-SOTOPIA, we find that goal achievement and believability of all of the language models that we test decline through the whole interaction. Although using an advanced memory method improves the agents' performance, the best agents still achieve a significantly lower goal completion rate than humans on scenarios requiring an explicit understanding of interaction history. These findings show that we can use LIFELONG-SOTOPIA to evaluate the social intelligence of language agents over lifelong social interactions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。