arXiv:2508.08501cs.AI2025-08被引 7

用无限生成的街机游戏测试大模型的推理与规划能力

GVGAI-LLM: Evaluating Large Language Model Agents with Infinite Games

  • 基于视频游戏语言构建可无限生成的新游戏基准
  • 118个游戏中大模型在空间推理上普遍出错
  • 适合研究智能体行为与空间认知的学者使用

我们提出GVGAI-LLM,一个基于通用视频游戏AI框架的视频游戏基准,用于评估大语言模型(LLMs)的推理与问题解决能力。该基准包含多样化的街机风格游戏,旨在测试模型处理不同于现有基准的任务的能力。利用视频游戏描述语言,可快速创建新游戏(含规则与关卡),防止模型过拟合。每个游戏场景由一组紧凑的ASCII字符表示,便于语言模型高效处理。GVGAI-LLM定义了可解释的评估指标,包括有意义步比、步效率和总分,以衡量模型表现。在118个具有多样化挑战与技能深度的游戏上进行零样本评估,揭示了当前模型在空间推理和基础规划上的持续局限性。模型普遍存在空间与逻辑错误,促使采用结构化提示与空间定位技术。尽管这些干预带来部分改进,该基准仍远未解决。GVGAI-LLM为推进语言模型能力研究提供了可复现的实验平台,尤其关注智能体行为与空间推理。其支持人工与程序化生成无限基准的能力,提供可扩展的长期评估框架。

原文摘要 · Abstract (English)

We introduce GVGAI-LLM, a video game benchmark for evaluating the reasoning and problem-solving capabilities of large language models (LLMs). Built on the General Video Game AI framework, it features a diverse collection of arcade-style games designed to test a model's ability to handle tasks that differ from most existing LLM benchmarks. The benchmark leverages a video game description language that enables the rapid creation of new games (including rules and levels), helping to prevent overfitting over time. Each game scene is represented by a compact set of ASCII characters, allowing for efficient processing by language models. GVGAI-LLM defines interpretable metrics, including meaningful step ratio, step efficiency, and overall score, to assess model behavior. Through zero-shot evaluations across 118 games with diverse challenges and skill depth, we reveal persistent limitations of LLMs in spatial reasoning and basic planning. Current models consistently exhibit spatial and logical errors, motivating structured prompting and spatial grounding techniques. Although these interventions lead to partial improvements, the benchmark remains very far from being solved. GVGAI-LLM serves as a reproducible testbed for advancing research on language model capabilities, with a particular emphasis on agentic behavior and spatial reasoning. Furthermore, its ability to generate infinite benchmarks, both manually and procedurally, provides a scalable framework for longitudinal evaluation.

大模型评估空间推理智能体行为游戏基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。