用20年足球俱乐部管理测试大模型长期决策能力
FM-Bench: A Benchmark for Long-Horizon Management with Competing Agents

- 让LLM代理在20年周期内管理足球俱乐部,完成400次决策
- 15个模型全部完成全程,最佳模型平均分领先但排名随时间变化
- 长期表现取决于管理行为而非算力,需合理规划投资与续约时机
语言模型代理已能可靠执行短期任务,但在长期决策场景中,其能否持续做出有效决策仍缺乏衡量。FM-Bench(足球管理基准)通过模拟20个游戏年、26个工具和约340至400个决策节点,评估大模型在复杂环境中的持久性决策能力。一个LLM代理需在与所有对手相同的预算下组建球队、交易球员、谈判合同、投资青训与设施、排定阵容,并应对董事会可能的解雇,由确定性引擎累计每赛季结果生成最终得分,无大模型评分或人工评判。独奏赛道将15个前沿模型置于固定脚本世界中进行对比,竞技场则让相同模型与脚本基线共同参与共享的20年世界,为首次此类规模的直接对抗评估。我们测量六项行为能力。在三个随机种子下,所有15个模型均完成整个周期,而盲选脚本基线大多中途失败;Claude-Fable-5在独奏中排名第一,在竞技场中也领先,但冠军名次在十模型间轮转。规模、价格或厂商均无法预测排名,排序仅在后期稳定,人类首播玩家反而位居模型末尾。真正区分模型的是管理行为:高分模型更少在后期投入长期回报项目,保持资金活跃而非闲置,并提前开启续约,而令牌消耗与表现无关。无模型能从数百次被拒报价中学习市场隐含价格,自管理记忆也失效于两种极端模式:不断积累的档案或每季重写的计划。代码已在https://github.com/Analogy-AI/fm-bench公开。
原文摘要 · Abstract (English)
Language model agents now execute bounded tasks reliably. Whether they can sustain effective decision-making over long horizons, where actions have cumulative consequences and the environment responds to their choices, remains largely unmeasured. FM-Bench (Football Management Benchmark) measures this. An LLM agent runs a football club for 20 in-game years through 26 tools and roughly 340 to 400 decision stops. It drafts a squad on the same budget as every rival, trades players, negotiates contracts, invests in facilities and youth, sets lineups, and answers to a board that can fire it, while a deterministic engine accumulates every year into one final score with no LLM judge or human rater. The solo track plays each of 15 frontier models against a frozen scripted world, and the Arena places the same models plus a scripted anchor in one shared 20-year world; to our knowledge, the first head-to-head evaluation at this scale. We measure six behavioral capabilities behind the score. Across three seeds, all 15 models complete every horizon while the blind scripted baselines die out in most of theirs, and claude-fable-5 tops the solo board on mean score and the Arena, where the title nonetheless rotates among ten models. Neither scale, price, nor vendor predicts the order; the order settles only late in the horizon, and the best first-play human lands only at the bottom of the model board. What separates the models is managerial behavior rather than computation. Higher-scoring models reduce slow-payoff investment near the end, keep cash invested rather than idle, and open renewals well before the deadline, while token spend predicts nothing. No model learns the market's hidden prices from hundreds of rejected bids, and self-managed memory fails in two opposite modes: an archive that only grows or a plan rewritten every season. Code is available at https://github.com/Analogy-AI/fm-bench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。