用游戏评测大模型多智能体系统,提升评估透明度与可复现性。
WiS Platform: Enhancing Evaluation of LLM-Based Multi-Agent Systems Through Game-Based Analysis
- 基于猜间谍游戏构建开放评测平台,支持Hugging Face模型接入
- 实时更新排行榜,覆盖胜率、攻防策略与推理能力综合评估
- 适合研究多智能体协作与大模型行为分析的学者使用
基于大语言模型(LLMs)的自主多智能体系统(MAS)在复杂任务处理中展现出强大能力,但现有研究在评估、分析和可复现性方面仍存在明显不足。本文提出WiS平台,通过《谁是间谍?》(Who is Spy?)游戏实现对LLM-based MAS的评测。平台具备三大特性:(1) 统一模型评估接口,支持Hugging Face上的多种模型;(2) 实时更新的排行榜;(3) 覆盖游戏胜率、攻击与防御策略、推理能力的全面评估。我们对多个开源与闭源LLM进行了广泛实验,发现不同智能体在游戏中的行为表现各异且富有洞察力。结果表明,该平台在评估效率与有效性方面具有显著优势。平台及文档已公开:https://whoisspy.ai/。
原文摘要 · Abstract (English)
Recent advancements in autonomous multi-agent systems (MAS) based on large language models (LLMs) have enhanced the application scenarios and improved the capability of LLMs to handle complex tasks. Despite demonstrating effectiveness, existing studies still evidently struggle to evaluate, analysis, and reproducibility of LLM-based MAS. In this paper, to facilitate the research on LLM-based MAS, we introduce an open, scalable, and real-time updated platform for accessing and analyzing the LLM-based MAS based on the games Who is Spy?" (WiS). Our platform is featured with three main worths: (1) a unified model evaluate interface that supports models available on Hugging Face; (2) real-time updated leaderboard for model evaluation; (3) a comprehensive evaluation covering game-winning rates, attacking, defense strategies, and reasoning of LLMs. To rigorously test WiS, we conduct extensive experiments coverage of various open- and closed-source LLMs, we find that different agents exhibit distinct and intriguing behaviors in the game. The experimental results demonstrate the effectiveness and efficiency of our platform in evaluating LLM-based MAS. Our platform and its documentation are publicly available at https://whoisspy.ai/.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。