评测大模型在英超赛季中的长期投注决策能力,发现现有模型普遍亏损。
KellyBench: A Benchmark for Long-Horizon Sequential Decision Making

- 构建英超2023-24赛季的序列化投注环境,模拟长期决策
- 所有前沿模型平均亏损8%,多模型遭遇破产风险
- 引入专家评分体系,揭示模型策略远逊人类水平
语言模型在目标明确的任务中已接近性能饱和,但正被部署于目标开放、环境动态的长期决策场景。本文提出KellyBench,一个用于评估体育博彩市场中序列决策能力的基准环境。模型需在2023-24赛季英格兰足球超级联赛的序列化仿真中,最大化长期资金增长。提供包括高级统计数据、首发名单和公众赔率在内的详尽历史数据。成功需构建机器学习模型、识别市场偏差并随环境变化调整策略。测试显示,所有评估的前沿模型在五组随机种子下均亏损,最优模型平均回报为-8%,许多模型在不同种子下遭遇破产。通过人类专家评分体系评估策略复杂度,发现模型表现远低于人类基线;Claude Opus 4.6得分为26.5%。该基准可通过 https://openreward.ai/GeneralReasoning/KellyBench 开放获取。
原文摘要 · Abstract (English)
Language models are saturating benchmarks for procedural tasks with narrow objectives. But they are increasingly being deployed in long-horizon, non-stationary environments with open-ended goals. In this paper we introduce KellyBench, an environment for evaluating sequential decision-making in sports betting markets. Agents are placed in a sequential simulation of the 2023-24 English Premier League season and tasked with maximising their long-term bankroll growth. They are given detailed historical data, including advanced statistics, lineups, and public odds. To succeed they must build machine learning models, identify edge in public markets, and adapt as the environment changes over time. We find that all frontier models evaluated lose money on average over the course of the season for five seeds. The best performing model achieves an average return of -8%, and many models experiencing ruin across seeds. To judge strategy sophistication, we use a human expert rubric to grade each model and find their approaches to be unsophisticated compared to human baselines; Claude Opus 4.6 achieves a rubric score of 26.5%, which means there is significant room for improvement. KellyBench is available as an open-access API endpoint at https://openreward.ai/GeneralReasoning/KellyBench.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。