arXiv:2502.20807cs.LG2025-02中稿 · ed被引 6

用策略游戏测试大模型像人一样玩游戏的能力

Digital Player: Evaluating Large Language Models based Human-like Agent in Games

  • 基于开源游戏Unciv构建测试平台,模拟真实决策与对话场景
  • 挑战大模型在复杂策略与长期规划中的推理与社交互动能力
  • 适合研究智能体交互、人机协作及大模型行为评估的学者

随着大型语言模型(LLMs)的快速发展,基于LLM的自主智能体展现出作为数字员工的潜力,例如数字分析师、教师和程序员。本文基于拥有数百万活跃玩家的开源策略游戏Unciv,构建了一个应用级测试平台,旨在帮助研究者建立“数据飞轮”,用于研究“数字玩家”任务中的人类类似智能体。这款类文明游戏具有广阔的决策空间以及丰富的语言交互,如外交谈判和欺骗行为,对基于LLM的智能体在数值推理和长期规划方面构成重大挑战。另一个挑战是让“数字玩家”生成符合人类特征的回应,以实现与人类玩家的社会互动、合作与协商。该项目开源,地址为 https://github.com/fuxiAIlab/CivAgent。

原文摘要 · Abstract (English)

With the rapid advancement of Large Language Models (LLMs), LLM-based autonomous agents have shown the potential to function as digital employees, such as digital analysts, teachers, and programmers. In this paper, we develop an application-level testbed based on the open-source strategy game "Unciv", which has millions of active players, to enable researchers to build a "data flywheel" for studying human-like agents in the "digital players" task. This "Civilization"-like game features expansive decision-making spaces along with rich linguistic interactions such as diplomatic negotiations and acts of deception, posing significant challenges for LLM-based agents in terms of numerical reasoning and long-term planning. Another challenge for "digital players" is to generate human-like responses for social interaction, collaboration, and negotiation with human players. The open-source project can be found at https:/github.com/fuxiAIlab/CivAgent.

大模型评估游戏智能体人机交互

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。