升级狼人杀评测平台,支持多角色自定义与多维度评估
WereWolf-Plus: An Update of Werewolf Game setting Based on DSGBench
- 构建可扩展的多角色狼人杀评测框架,支持自定义角色配置
- 引入针对各角色的量化评估指标,覆盖推理、合作与社交影响力
- 适合研究多智能体策略推理与社会交互的学者使用
随着基于大语言模型的智能体快速发展,其社交互动与策略推理能力日益受到关注。然而现有基于狼人杀的评测平台存在游戏设定过于简单、评估指标不完整及可扩展性差等问题。为此,我们提出WereWolf-Plus,一个面向多智能体策略推理的多模型、多维度、多方法评测平台。该平台具备强可扩展性,支持先知、女巫、猎人、守卫、警长等角色的自定义配置,以及不同角色的模型分配与推理增强策略。此外,我们设计了一套全面的定量评估指标,涵盖所有特殊角色、狼人及警长,并丰富了对智能体推理能力、协作能力与社会影响力的评估维度。WereWolf-Plus为多智能体社区中的推理与策略交互研究提供了更灵活可靠的评估环境。代码已开源:https://github.com/MinstrelsyXia/WereWolfPlus。
原文摘要 · Abstract (English)
With the rapid development of LLM-based agents, increasing attention has been given to their social interaction and strategic reasoning capabilities. However, existing Werewolf-based benchmarking platforms suffer from overly simplified game settings, incomplete evaluation metrics, and poor scalability. To address these limitations, we propose WereWolf-Plus, a multi-model, multi-dimensional, and multi-method benchmarking platform for evaluating multi-agent strategic reasoning in the Werewolf game. The platform offers strong extensibility, supporting customizable configurations for roles such as Seer, Witch, Hunter, Guard, and Sheriff, along with flexible model assignment and reasoning enhancement strategies for different roles. In addition, we introduce a comprehensive set of quantitative evaluation metrics for all special roles, werewolves, and the sheriff, and enrich the assessment dimensions for agent reasoning ability, cooperation capacity, and social influence. WereWolf-Plus provides a more flexible and reliable environment for advancing research on inference and strategic interaction within multi-agent communities. Our code is open sourced at https://github.com/MinstrelsyXia/WereWolfPlus.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。