评测大模型在多人交互环境中的规划与社交推理能力
SocialGrid: A Benchmark for Planning and Social Reasoning in Embodied Multi-Agent Systems

- 基于《 Among Us 》设计多智能体环境,评估规划与社交推理
- 最强模型任务完成率低于60%,导航问题干扰社交评估
- 提供计划辅助与失败分析工具,适合智能体开发调试
随着大型语言模型(LLMs)从文本处理转向自主代理,评估其在具身多智能体场景下的社交推理能力变得至关重要。我们提出 SocialGrid,一个受《Among Us》启发的具身多智能体环境,用于评估 LLM 代理在规划、任务执行和社交推理方面的能力。评估显示,即使最强的开源模型(GPT-OSS-120B)在任务完成与规划上的准确率也低于60%,智能体常陷入重复行为或无法通过基本障碍。由于糟糕的导航会干扰社交智能评估,SocialGrid 提供可选的 Planning Oracle 来分离社交推理与规划缺陷。尽管规划辅助提升了任务完成率,社交推理仍为瓶颈:无论规模大小,智能体检测欺骗的准确率接近随机水平,仅依赖浅层启发式而非累积行为证据。SocialGrid 提供自动故障分析与细粒度指标,帮助开发者诊断并改进智能体。我们还通过对抗性联赛赛制建立基于 Elo 评分的竞争排行榜。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) transition from text processors to autonomous agents, evaluating their social reasoning in embodied multi-agent settings becomes critical. We introduce SocialGrid, an embodied multi-agent environment inspired by Among Us that evaluates LLM agents on planning, task execution, and social reasoning. Our evaluations reveal that even the strongest open model (GPT-OSS-120B) achieves below 60% accuracy in task completion and planning, with agents getting stuck in repetitive behaviors or failing to navigate basic obstacles. Since poor navigation confounds evaluation of social intelligence, SocialGrid offers an optional Planning Oracle to isolate social reasoning from planning deficits. While planning assistance improves task completion, social reasoning remains a bottleneck: agents fail to detect deception at near-random chance regardless of scale, relying on shallow heuristics rather than accumulating behavioral evidence. SocialGrid provides automatic failure analysis and fine-grained metrics, enabling developers to diagnose and improve their agents. We also establish a competitive leaderboard using Elo ratings from adversarial league play.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。