arXiv:2510.26298cs.CLcs.AI2025-10被引 1

测试ChatGPT Atlas在网页游戏中的表现,发现它逻辑强但反应慢。

Can Agent Conquer Web? Exploring the Frontiers of ChatGPT Atlas Agent in Web Games

  • 用网页游戏作为测试场景,评估模型的真实交互能力。
  • 解数独快于人类,但在需要精准操作的游戏中难以通过初始关卡。
  • 适合关注AI网页交互局限性的研究者与开发者参考。

OpenAI的ChatGPT Atlas引入了网页交互新能力,可分析页面、理解用户意图,并直接在浏览器中执行鼠标和键盘操作。尽管其信息检索能力已被验证,但在动态交互环境中的表现仍不明确。本研究以基于浏览器的游戏(包括Google T-Rex Runner、数独、Flappy Bird和Stein.world)为测试场景,采用游戏得分作为量化指标,评估不同任务类型下的表现。结果表明,Atlas在逻辑推理类任务(如数独)中表现优异,完成速度显著快于人类基准;但在需要精确时序与运动控制的实时游戏中,常无法越过初始障碍。这说明尽管其分析处理能力较强,但在需实时交互的动态网页环境中仍存在明显瓶颈。项目主页:https://atlas-game-eval.github.io。

原文摘要 · Abstract (English)

OpenAI's ChatGPT Atlas introduces new capabilities for web interaction, enabling the model to analyze webpages, process user intents, and execute cursor and keyboard inputs directly within the browser. While its capacity for information retrieval tasks has been demonstrated, its performance in dynamic, interactive environments remains less explored. In this study, we conduct an early evaluation of Atlas's web interaction capabilities using browser-based games as test scenarios, including Google's T-Rex Runner, Sudoku, Flappy Bird, and Stein.world. We employ in-game performance scores as quantitative metrics to assess performance across different task types. Our results show that Atlas performs strongly in logical reasoning tasks like Sudoku, completing puzzles significantly faster than human baselines, but struggles substantially in real-time games requiring precise timing and motor control, often failing to progress beyond initial obstacles. These findings suggest that while Atlas demonstrates capable analytical processing, there remain notable limitations in dynamic web environments requiring real-time interaction. The website of our project can be found at https://atlas-game-eval.github.io.

AI代理网页交互游戏测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。