arXiv:2604.07733cs.AI2026-04被引 2

用《文明5》测试大模型战略决策,让评估更精细。

CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V

  • 用每回合状态预测胜率,替代仅看输赢
  • 307场对局验证模型真实战略能力
  • 适合研究智能体策略差异与长期决策

评估基于大模型的智能体在战略决策中的表现,需要生成性、竞争性和长期性的环境。然而,现有基准大多缺乏三者之一,且评估信号不足以支持长周期多智能体对战。我们提出CivBench,一个用于多人《文明5》游戏中大模型策略智能体的评测基准。由于游戏持续数百回合、面对多个对手,最终胜负信号过于稀疏,CivBench通过训练模型基于回合级游戏状态估计胜率,并通过预测效度、构念效度和收敛效度进行验证。在307场包含7个大模型及多种代理设置的对局中,我们展示了CivBench作为非饱和基准在估计战略能力方面的潜力,揭示了不同智能体设置带来的模型特异性影响,并识别出仅靠结果无法观察到的差异化战略特征。

原文摘要 · Abstract (English)

Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offer evaluation signals rich enough for long-horizon, multi-agent play. We introduce CivBench, a benchmark for LLM strategists (i.e., agentic setups) in multiplayer Civilization V. Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains models on turn-level game state to estimate victory probabilities throughout play, validated through predictive, construct, and convergent validity. Across 307 games with 7 LLMs and multiple CivBench agent conditions, we demonstrate CivBench's potential to estimate strategic capabilities as an unsaturated benchmark, reveal model-specific effects of agentic setup, and outline distinct strategic profiles not visible through outcome-only evaluation.

战略决策大模型评测游戏智能体长期规划

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。