arXiv:2504.11442cs.CLcs.AI2025-04被引 7

用文字游戏评测大模型的社交智能,支持多人实时对战。

TextArena

  • 构建57个以上文本游戏环境,支持单人、双人及多人对战。
  • 通过实时TrueSkill评分系统,可在线与人类或模型对战评估。
  • 专为研究社区设计,易扩展新游戏和训练模型,适合智能体研究者。

TextArena 是一个开源的、基于文本的竞争性游戏集合,用于训练和评估大型语言模型(LLMs)的智能体行为。它涵盖57个以上的独特环境,包括单人、双人和多人设置,并通过在线对战系统实现模型能力的实时评估(可对抗人类和其他提交的模型),并提供实时的TrueSkill评分。传统基准测试很少评估动态社交技能,如谈判、心智理论和欺骗,而TextArena填补了这一空白。该平台以研究、社区和可扩展性为核心设计,强调新增游戏、框架调整、模型测试、对战和训练的便捷性。详细的游戏环境、规则、排行榜和示例文档可在 https://github.com/LeonGuertler/TextArena 和 https://www.textarena.ai/ 获取。

原文摘要 · Abstract (English)

TextArena is an open-source collection of competitive text-based games for training and evaluation of agentic behavior in Large Language Models (LLMs). It spans 57+ unique environments (including single-player, two-player, and multi-player setups) and allows for easy evaluation of model capabilities via an online-play system (against humans and other submitted models) with real-time TrueSkill scores. Traditional benchmarks rarely assess dynamic social skills such as negotiation, theory of mind, and deception, creating a gap that TextArena addresses. Designed with research, community and extensibility in mind, TextArena emphasizes ease of adding new games, adapting the framework, testing models, playing against the models, and training models. Detailed documentation of environments, games, leaderboard, and examples are available on https://github.com/LeonGuertler/TextArena and https://www.textarena.ai/.

智能体评估文本游戏社交智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。