arXiv:2604.08340cs.CVcs.AI2026-04被引 1

提出可跨回合自适应的多模态强化学习框架,解决视觉游戏中的策略演化难题。

Mastering PokeGym: Graph-Guided Multimodal Evolution at Test Time

  • 用图结构协同优化视觉、策略与动作,实现多模态联合进化
  • 在PokeGym上达成60.18%成功率,比最强基线高11个百分点
  • 适合研究测试时学习与多模态协同优化的学者参考

尽管人工智能已在棋类等结构化游戏中取得突破,但视觉驱动的3D游戏中,视觉-语言代理仍因无法访问游戏状态而表现不佳。现有环境通常评估固定配置的代理,而非其在连续任务中改进自身配置的能力——即测试时学习(TTL)。此外,当前TTL方法多孤立优化单一模态(如文本提示或动作),忽视感知、推理与控制间的协同效应。为此,我们首先构建了基于《宝可梦传说:阿尔宙斯》的长周期基准测试集PokeGym,要求代理仅凭视觉观测完成任务,评估其跨回合学习与适应能力。为应对该挑战,提出图引导的多模态协同演化框架G-EvoMAC,联合优化视觉感知、策略与动作集。大量实验表明,G-EvoMAC在PokeGym上达到60.18%平均成功率,超越最强基线超11个百分点,验证了跨模态协同演化的有效性。

原文摘要 · Abstract (English)

While artificial intelligence has mastered structured games like chess and Go, vision-language agents still struggle in visually-driven 3D games without access to game states. Existing game environments typically evaluate a fixed agent configuration, rather than an agent's ability to improve its configuration across consecutive episodes of the same task---a paradigm known as test-time learning (TTL). Furthermore, current TTL methods typically optimize single modalities---such as text prompts or actions---in isolation, ignoring the synergy between perception, reasoning, and control. To bridge these gaps, we first introduce \textbf{PokeGym}, a long-horizon benchmark built upon the 3D open-world game Pokémon Legends: Z-A, where agents act from visual observations without access to game states, designed to evaluate an agent's ability to learn and adapt across consecutive episodes of the task. To tackle this challenging environment, we propose Graph-Guided Evolutionary Multimodal Agent Configuration (\textbf{G-EvoMAC}), a graph-guided framework that jointly optimizes visual perception, strategy, and action set synergistically. Extensive experiments show that G-EvoMAC achieves a 60.18\% average success rate on PokeGym, outperforming the strongest baseline by over 11 percentage points, validating the power of cross-modal co-evolution.

测试时学习多模态协同视觉代理强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。