arXiv:2509.00483cs.CV2025-09

用跳跳游戏测试大模型决策能力,发现其空间判断有短板

Exploring Decision-Making Capabilities of LLM Agents: An Experimental Study on Jump-Jump Game

论文配图:Exploring Decision-Making Capabilities of LLM Agents: An Experimental Study on Jump-Jump Game
图 1 · 摘自论文原文
  • 用跳跳游戏测试大模型的跳跃力控制与路径规划能力
  • 模型在复杂场景下准确率不足60%,依赖简单规则
  • 适合研究具身智能与决策模型的开发者参考

跳跳游戏作为一种简单但具有挑战性的休闲游戏,为研究大语言模型(LLM)的决策能力提供了理想的测试环境。游戏中玩家需根据当前位置与目标平台的距离,精确控制跳跃力度,涉及空间推理、物理建模和策略规划等多个认知层面。游戏机制描述了玩家角色(红色圆圈)必须以合适的跳跃力跨越平台以最大化得分的过程。该实验旨在评估大模型在动态环境下的实时决策表现,揭示其在复杂情境中执行连续动作的局限性。

原文摘要 · Abstract (English)

The Jump-Jump game, as a simple yet challenging casual game, provides an ideal testing environment for studying LLM decision-making capabilities. The game requires players to precisely control jumping force based on current position and target platform distance, involving multiple cognitive aspects including spatial reasoning, physical modeling, and strategic planning. It illustrates the basic gameplay mechanics of the Jump-Jump game, where the player character (red circle) must jump across platforms with appropriate force to maximize score.

大模型决策游戏测试空间推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。