用游戏锦标赛评测大模型社交能力,还能自动生成改进策略。
Social Gym and SPaRTan: Benchmarking and Improving LLM Social Reasoning via Multi-Agent Game Tournaments

- 构建21种多人社交游戏环境,用规则判定结果实现客观评估。
- GPT-5-mini表现最佳,但无模型在所有角色中都优秀。
- 提出SPaRTan自反思机制,可生成可复用的策略手册提升表现。
大语言模型代理正越来越多地应用于多代理社交场景,需具备合作、协商和适应能力。但社交技能难以评估,因缺乏客观标准,现有评价依赖成本高、主观性强的LLM裁判,且模型无法获得可靠学习信号。为此,我们提出Social Gym——一个包含21种多代理社交游戏(如狼人杀、抵抗组织、猜间谍)的环境,其规则决定的结果使代理表现可验证、客观,并通过埃洛锦标赛生成跨游戏排名。基准测试显示,尽管GPT-5-mini位居榜首,但无模型在所有游戏中均表现优异或在所有角色中胜出,暴露出社交推理的局限性。受此启发,我们进一步提出SPaRTan(自对弈与反思迁移):一种无需权重更新的自我改进循环——模型先玩游戏,反思轨迹与结果生成可迁移的策略手册,并在后续游戏中应用。实验表明,SPaRTan策略手册能帮助GPT-5-mini在较弱角色中提升表现,但对Qwen3-32B提升有限。Social Gym与SPaRTan共同提供了无需权重更新的可复现、可验证的大模型社交推理评测与改进基础。
原文摘要 · Abstract (English)
LLM agents are increasingly deployed in multi-agent social settings where they must cooperate, negotiate, and adapt to other agents. Measuring and improving these social skills is hard because, unlike math or logic, social interaction offers no objective ground truth: evaluations fall back on LLM judges, which are costly, subjective, and noisy, and models get no reliable signal to learn from. To address both, we first introduce Social Gym, an environment of 21 multi-agent social games (e.g., Werewolves, Resistance, Spyfall) whose rule-decided outcomes make agent performance verifiable and objective, with an Elo tournament that produces a cross-game leaderboard. Benchmarking experiments show that while GPT-5-mini tops the leaderboard, no model excels at all games uniformly or in all game roles, pointing to limitations of social reasoning. Motivated by this, we additionally propose SPaRTan (Self-Play and Reflect-Transfer), a training-free self-improvement loop: a model plays a game, reflects on its trajectories and their outcomes to produce a transferable playbook, and applies that playbook in subsequent games. Our results show that SPaRTan playbooks help GPT-5-mini agents level their performance on weaker roles, but largely do not improve Qwen3-32B's performance. Together, Social Gym and SPaRTan offer a reproducible, verifiable foundation for measuring and improving LLM social reasoning without weight updates.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。