用《文明5》测试大模型战略决策,让评估更精细。
CivBench: Progress-Based Evaluation for LLMs' Strategic Decision-Making in Civilization V
- 用每回合状态预测胜率,替代仅看输赢
- 307场对局验证模型真实战略能力
- 适合研究智能体策略差异与长期决策
评估基于大模型的智能体在战略决策中的表现,需要生成性、竞争性和长期性的环境。然而,现有基准大多缺乏三者之一,且评估信号不足以支持长周期多智能体对战。我们提出CivBench,一个用于多人《文明5》游戏中大模型策略智能体的评测基准。由于游戏持续数百回合、面对多个对手,最终胜负信号过于稀疏,CivBench通过训练模型基于回合级游戏状态估计胜率,并通过预测效度、构念效度和收敛效度进行验证。在307场包含7个大模型及多种代理设置的对局中,我们展示了CivBench作为非饱和基准在估计战略能力方面的潜力,揭示了不同智能体设置带来的模型特异性影响,并识别出仅靠结果无法观察到的差异化战略特征。
原文摘要 · Abstract (English)
Evaluating strategic decision-making in LLM-based agents requires generative, competitive, and longitudinal environments, yet few benchmarks provide all three, and fewer still offer evaluation signals rich enough for long-horizon, multi-agent play. We introduce CivBench, a benchmark for LLM strategists (i.e., agentic setups) in multiplayer Civilization V. Because terminal win/loss is too sparse a signal in games spanning hundreds of turns and multiple opponents, CivBench trains models on turn-level game state to estimate victory probabilities throughout play, validated through predictive, construct, and convergent validity. Across 307 games with 7 LLMs and multiple CivBench agent conditions, we demonstrate CivBench's potential to estimate strategic capabilities as an unsaturated benchmark, reveal model-specific effects of agentic setup, and outline distinct strategic profiles not visible through outcome-only evaluation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。