构建游戏基准测试,评估智能体完成完整剧情的能力。
FlashAdventure: A Benchmark for GUI Agents Solving Full Story Arcs in Diverse Adventure Games
- 设计34款闪存冒险游戏,检验智能体全程解谜能力
- 提出新框架使智能体能记住早期线索,提升任务完成率
- 适合研究人机交互与长程规划的学者参考
基于大模型的GUI智能体在互动数字环境方面展现出潜力。其中,电子游戏因其界面多样,成为理想的测试平台,而解谜类游戏因叙事驱动的复杂交互带来额外挑战。现有游戏基准缺乏多样性,且很少评估智能体完成完整故事线的能力。为此,我们提出FlashAdventure,一个包含34款基于Flash的冒险游戏的基准,用于测试智能体完成全剧情弧线的能力,并解决观察-行为鸿沟问题:即记忆并运用早期游戏信息的难题。我们还提出了CUA-as-a-Judge自动化评估工具,以及COAST智能体框架,通过长期线索记忆实现更优的序列任务规划与求解。实验表明,当前GUI智能体难以完成完整剧情,而COAST通过弥合观察-行为鸿沟显著提升了里程碑完成率。然而,人类与最优智能体之间仍存在显著差距,需持续投入研究以缩小这一鸿沟。
原文摘要 · Abstract (English)
GUI agents powered by LLMs show promise in interacting with diverse digital environments. Among these, video games offer a valuable testbed due to their varied interfaces, with adventure games posing additional challenges through complex, narrative-driven interactions. Existing game benchmarks, however, lack diversity and rarely evaluate agents on completing entire storylines. To address this, we introduce FlashAdventure, a benchmark of 34 Flash-based adventure games designed to test full story arc completion and tackle the observation-behavior gap: the challenge of remembering and acting on earlier gameplay information. We also propose CUA-as-a-Judge, an automated gameplay evaluator, and COAST, an agentic framework leveraging long-term clue memory to better plan and solve sequential tasks. Experiments show current GUI agents struggle with full story arcs, while COAST improves milestone completion by bridging the observation-behavior gap. Nonetheless, a marked discrepancy between humans and best-performing agents warrants continued research efforts to narrow this divide.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。