arXiv:2410.02829cs.AIcs.HC2024-10被引 7

用大模型当游戏难度测试员,效果比想象中好。

LLMs May Not Be Human-Level Players, But They Can Be Testers: Measuring Game Difficulty with LLM Agents

  • 用提示词引导大模型玩策略游戏,自动评估难度
  • 模型表现与人类玩家评价高度相关,相关性显著
  • 适合游戏开发团队快速验证关卡难度

大语言模型在多项任务中展现出作为自主智能体的潜力。本文探索游戏行业中一个实际问题:大模型能否用于衡量游戏难度?我们提出一种通用的游戏测试框架,利用大模型代理,在两款广受欢迎的策略类游戏——Wordle 和 Slay the Spire 上进行测试。结果显示,尽管大模型表现不如普通人类玩家,但通过简单的通用提示技术引导后,其评估结果与人类玩家对难度的判断存在统计上显著且强相关的关联。这表明大模型可作为游戏开发过程中有效的难度测评工具。基于实验结果,我们进一步总结了将大模型融入游戏测试流程的一般原则与建议。

原文摘要 · Abstract (English)

Recent advances in Large Language Models (LLMs) have demonstrated their potential as autonomous agents across various tasks. One emerging application is the use of LLMs in playing games. In this work, we explore a practical problem for the gaming industry: Can LLMs be used to measure game difficulty? We propose a general game-testing framework using LLM agents and test it on two widely played strategy games: Wordle and Slay the Spire. Our results reveal an interesting finding: although LLMs may not perform as well as the average human player, their performance, when guided by simple, generic prompting techniques, shows a statistically significant and strong correlation with difficulty indicated by human players. This suggests that LLMs could serve as effective agents for measuring game difficulty during the development process. Based on our experiments, we also outline general principles and guidelines for incorporating LLMs into the game testing process.

游戏测试大模型应用难度评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。