arXiv:2605.29653cs.AI2026-05

用宝可梦卡牌游戏测试大模型代理的策略学习与自我进化能力

PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?

论文配图:PTCG-Bench: Can LLM Agents Master Pokémon Trading Card Game?
图 1 · 摘自论文原文
  • 基于宝可梦卡牌游戏构建双层评估框架,测策略决策与自进化
  • 大模型代理表现尚可但持续稳定进化仍困难,结果受评测设计影响大
  • 适合研究自适应智能体、交互式环境中的自主学习者

面对策略复杂的棋类游戏,人类玩家在几轮对局后即可快速制定策略。而自主智能体需具备类似能力,但现有基准常无法充分捕捉此类动态决策场景。我们提出PTCG-Bench,基于宝可梦卡牌游戏(PTCG)构建的基准,从两个互补层面评估大模型代理:(1) 单一复杂环境下的决策表现;(2) 通过积累经验实现自我演化的能力。同时引入模块化评测框架消融实验,以更清晰地解析性能差异,避免与模型能力混淆。实验表明,尽管大模型代理能实现非平凡的游戏表现,但持续且稳定的自我进化仍具挑战,且性能对评测设计敏感。我们希望PTCG-Bench能推动未来在真实交互环境中注重评测设计与自我演化的智能体研究。

原文摘要 · Abstract (English)

Given a strategically complex board game, human players can quickly learn to devise strategies after playing a few rounds. Autonomous agents require similar capabilities in realistic interactive environments, yet existing agent benchmarks often fail to fully capture such strategic and evolving decision-making scenarios. We present PTCG-Bench, a benchmark built on the Pok'{e}mon Trading Card Game (PTCG) that evaluates LLM agents at two complementary levels: (1) their decision-making performance within a single complex environment, and (2) their ability to self-evolving through accumulated experience. We further include a modular harness ablation to better interpret agent performance without conflating it with model capability. Our experiments show that, although LLM agents can achieve non-trivial gameplay performance, sustained and stable self-evolution remains challenging, and performance is sensitive to harness design. We hope that PTCG-Bench will facilitate future research on harness-aware and self-evolving agents in realistic interactive environments.

大模型代理自进化游戏智能

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。