用固定等级的AI锚点实现大模型策略能力的稳定评估
BotzoneBench: Scalable LLM Evaluation via Graded AI Anchors
- 用预设技能等级的AI作为基准,替代动态对战评分
- 在8种游戏中测试5个模型,覆盖17万+状态动作对
- 可跨时间对比,适合长期追踪交互式AI性能
大型语言模型在需要战略决策的交互环境中日益普及,但系统性评估其能力仍具挑战。现有基准多聚焦静态推理任务,难以捕捉动态策略能力。近期基于游戏的评估采用模型间对战,依赖临时模型池,存在二次计算开销且缺乏稳定性能锚点。本文提出将评估锚定于技能校准的固定层级游戏AI,实现线性时间的绝对技能度量与跨时稳定的可解释性。基于Botzone平台,在涵盖确定性完美信息棋类到随机不完美信息卡牌游戏的八类游戏中,对五个主流模型进行系统评估,分析177,047个状态-动作对,揭示显著性能差异并识别不同战略行为,顶尖模型在多个领域表现接近中高阶专用游戏AI。该锚定范式可推广至任意具备明确技能层级的领域,构建可扩展、可复用的交互式AI评估框架。
原文摘要 · Abstract (English)
Large Language Models (LLMs) are increasingly deployed in interactive environments requiring strategic decision-making, yet systematic evaluation of these capabilities remains challenging. Existing benchmarks for LLMs primarily assess static reasoning through isolated tasks and fail to capture dynamic strategic abilities. Recent game-based evaluations employ LLM-vs-LLM tournaments that produce relative rankings dependent on transient model pools, incurring quadratic computational costs and lacking stable performance anchors for longitudinal tracking. The central challenge is establishing a scalable evaluation framework that measures LLM strategic reasoning against consistent, interpretable standards rather than volatile peer models. Here we show that anchoring LLM evaluation to fixed hierarchies of skill-calibrated game Artificial Intelligence (AI) enables linear-time absolute skill measurement with stable cross-temporal interpretability. Built on the Botzone platform's established competitive infrastructure, our BotzoneBench evaluates LLMs across eight diverse games spanning deterministic perfect-information board games to stochastic imperfect-information card games. Through systematic assessment of 177,047 state-action pairs from five flagship models, we reveal significant performance disparities and identify distinct strategic behaviors, with top-performing models achieving proficiency comparable to mid-to-high-tier specialized game AI in multiple domains. This anchored evaluation paradigm generalizes beyond games to any domain with well-defined skill hierarchies, establishing a scalable and reusable framework for assessing interactive AI capabilities.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。