用宝可梦对战测试大模型的战略推理能力
A Multi-Agent Pokemon Tournament for Evaluating Strategic Reasoning of Large Language Models
- 让大模型扮演训练家,参与宝可梦对战比赛
- 记录战术选择与换宠决策,分析策略深度
- 适合研究大模型博弈与决策能力的学者
本研究提出LLM宝可梦联盟,一个基于大型语言模型(LLMs)的智能体竞赛系统,用于模拟宝可梦对战中的战略决策。该平台在类型相克、回合制战斗环境中,评估不同LLMs在策略性、适应性与战术深度方面的表现。通过单淘汰赛制,系统记录详尽的决策日志,包括组队逻辑、行动选择与换宠策略。研究揭示现代大模型在不确定性下的理解、适应与优化能力,使宝可梦联盟成为评估人工智能战略推理与竞争学习的新基准。
原文摘要 · Abstract (English)
This research presents LLM Pokemon League, a competitive tournament system that leverages Large Language Models (LLMs) as intelligent agents to simulate strategic decision-making in Pokémon battles. The platform is designed to analyze and compare the reasoning, adaptability, and tactical depth exhibited by different LLMs in a type-based, turn-based combat environment. By structuring the competition as a single-elimination tournament involving diverse AI trainers, the system captures detailed decision logs, including team-building rationale, action selection strategies, and switching decisions. The project enables rich exploration into comparative AI behavior, battle psychology, and meta-strategy development in constrained, rule-based game environments. Through this system, we investigate how modern LLMs understand, adapt, and optimize decisions under uncertainty, making Pokémon League a novel benchmark for AI research in strategic reasoning and competitive learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。