arXiv:2602.17594cs.AI2026-02被引 8

用人类设计的无限游戏评估AI通用智能,发现当前模型表现远低于人类。

AI Gamestore: Scalable, Open-Ended Evaluation of Machine General Intelligence with Human Games

  • 通过大模型与人类协作生成海量人类游戏,构建开放评估平台。
  • 7个前沿视觉语言模型在100款新游戏中平均得分不足人类10%。
  • 特别在世界模型、记忆和规划类游戏中表现差,揭示AI短板。

在技术快速发展的背景下,对机器智能进行广义人类智能的严格评估变得日益重要且困难。传统基准仅衡量特定狭窄能力,且多为静态,易被开发者优化饱和。我们提出,更有效的评估方式是通过高度通用的游戏玩法:研究AI在与人类玩家同等经验、时间等资源条件下,学习并玩尽所有可想象的人类游戏的能力。我们将'人类游戏'定义为由人设计给人玩的游戏,并主张其构成的'人类游戏多元宇宙'具有良好的评估潜力。为此,我们引入了AI GameStore,一个可扩展的开放式平台,利用大模型与人类协同,从主流数字游戏平台自动获取并适配标准化、容器化游戏环境,生成代表性人类游戏。作为概念验证,我们基于苹果应用商店和Steam热门榜单生成了100款游戏,评估了7个前沿视觉语言模型在短回合游戏中的表现。结果显示,最佳模型在多数游戏中得分低于人类平均分的10%,尤其在考验世界模型、记忆与规划能力的游戏上表现不佳。最后,我们提出下一步建设方向,推动AI GameStore成为衡量和促进机器向人类级通用智能演进的实用工具。

原文摘要 · Abstract (English)

Rigorously evaluating machine intelligence against the broad spectrum of human general intelligence has become increasingly important and challenging in this era of rapid technological advance. Conventional AI benchmarks typically assess only narrow capabilities in a limited range of human activity. Most are also static, quickly saturating as developers explicitly or implicitly optimize for them. We propose that a more promising way to evaluate human-like general intelligence in AI systems is through a particularly strong form of general game playing: studying how and how well they play and learn to play \textbf{all conceivable human games}, in comparison to human players with the same level of experience, time, or other resources. We define a "human game" to be a game designed by humans for humans, and argue for the evaluative suitability of this space of all such games people can imagine and enjoy -- the "Multiverse of Human Games". Taking a first step towards this vision, we introduce the AI GameStore, a scalable and open-ended platform that uses LLMs with humans-in-the-loop to synthesize new representative human games, by automatically sourcing and adapting standardized and containerized variants of game environments from popular human digital gaming platforms. As a proof of concept, we generated 100 such games based on the top charts of Apple App Store and Steam, and evaluated seven frontier vision-language models (VLMs) on short episodes of play. The best models achieved less than 10\% of the human average score on the majority of the games, and especially struggled with games that challenge world-model learning, memory and planning. We conclude with a set of next steps for building out the AI GameStore as a practical way to measure and drive progress toward human-like general intelligence in machines.

通用智能游戏评估大模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。