arXiv:2603.15563cs.LGcs.AI2026-03被引 11

用宝可梦游戏构建大规模决策挑战,测试AI在复杂环境中的长期规划与博弈能力。

The PokeAgent Challenge: Competitive and Long-Context Learning at Scale

  • 基于宝可梦对战与RPG系统设计双赛道,融合部分可观知、博弈推理与长程规划。
  • 提供超2000万场对战轨迹数据,支持强化学习与大模型基线对比评估。
  • 适合研究长序列决策、多智能体协作及大模型泛化能力的学者与开发者。

我们提出PokeAgent挑战,一个基于宝可梦多智能体对战系统和庞大角色扮演环境的大规模决策研究基准。部分可观知、博弈论推理和长时程规划仍是前沿AI的开放问题,但现有基准极少能同时在真实条件下强调三者。PokeAgent通过两个互补赛道实现突破:对战赛道要求在部分可观知下进行战略推理与泛化,速度通关赛道则考验长时程规划与顺序决策。对战赛道提供超过2000万条对战轨迹数据集,配套启发式、强化学习与大模型基线,可实现高水平竞技表现。速度通关赛道首次建立标准化评估框架,包含开源多智能体协同系统,支持基于提示的大型语言模型方法模块化、可复现比较。2025年NeurIPS竞赛验证了资源质量与社区关注度,超100支队伍参与两赛道,优胜方案详述于论文中。参赛结果与基线分析揭示通用型(大模型)、专业型(强化学习)与顶尖人类表现间存在显著差距。基于BenchPress评估矩阵的分析表明,宝可梦对战几乎与标准大模型基准正交,测量现有套件未覆盖的能力,使宝可梦成为亟待突破的研究基准。我们推出动态更新的线上排行榜(对战)与自包含评估系统(速度通关),网址为https://pokeagentchallenge.com。

原文摘要 · Abstract (English)

We present the PokeAgent Challenge, a large-scale benchmark for decision-making research built on Pokemon's multi-agent battle system and expansive role-playing game (RPG) environment. Partial observability, game-theoretic reasoning, and long-horizon planning remain open problems for frontier AI, yet few benchmarks stress all three simultaneously under realistic conditions. PokeAgent targets these limitations at scale through two complementary tracks: our Battling Track, which calls for strategic reasoning and generalization under partial observability in competitive Pokemon battles, and our Speedrunning Track, which requires long-horizon planning and sequential decision-making in the Pokemon RPG. Our Battling Track supplies a dataset of 20M+ battle trajectories alongside a suite of heuristic, RL, and LLM-based baselines capable of high-level competitive play. Our Speedrunning Track provides the first standardized evaluation framework for RPG speedrunning, including an open-source multi-agent orchestration system for modular, reproducible comparisons of harness-based LLM approaches. Our NeurIPS 2025 competition validates both the quality of our resources and the research community's interest in Pokemon, with over 100 teams competing across both tracks and winning solutions detailed in our paper. Participant submissions and our baselines reveal considerable gaps between generalist (LLM), specialist (RL), and elite human performance. Analysis against the BenchPress evaluation matrix shows that Pokemon battling is nearly orthogonal to standard LLM benchmarks, measuring capabilities not captured by existing suites and positioning Pokemon as an unsolved benchmark that can drive RL and LLM research forward. We transition to a living benchmark with a live leaderboard for Battling and self-contained evaluation for Speedrunning at https://pokeagentchallenge.com.

决策智能多智能体长序列大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。