用世界杯比赛实时测试大模型预测能力,看它能否真预判未来。
LLM-SoccerArena: Benchmarking LLMs on Real-World Predictions in Sports

- 构建实时动态评测平台,让大模型在比赛未结束时作预测。
- 7个大模型预测104场世界杯赛,有网络访问的模型略胜一筹(Brier得分提升0.023)。
- 适合关注模型真实世界决策能力的研究者和开发者。
大型语言模型(LLMs)越来越多地用于支持对不确定未来事件的决策,但评估其预测真实世界结果的能力仍具挑战性。现有基准多为静态回溯式,无法检验模型在不确定性下合成信息预测未来的能力。我们提出 LLM-SoccerArena(https://llm-soccerarena.com),一个前瞻性实时评测基准,用于评估大模型在结果未知前对体育赛事的预测表现。该平台提供:(1) 前瞻性实时评测协议,(2) 公开开源平台,(3) 因子设计与赛事相关问题(如哪队获胜)。系统自动记录未决事件的时间戳、结构化验证的预测,以及提示、模型版本、工具调用轨迹与成本。因子设计涵盖四个维度:(1) 模型版本(如 GPT-5.5、Claude Opus 4.8),(2) 信息获取,(3) 提示策略,(4) 预测时间范围。我们在2026年FIFA世界杯上进行大规模评估,7个大模型对全部104场比赛及15个赛事相关问题生成预测。分析显示,具备网络访问能力的模型表现更优,但优势微弱(Brier得分提升0.023)。总体而言,LLM-SoccerArena 提供了一个灵活、开源的前瞻评测平台,将持续更新,适用于未来国内外赛事与联赛。
原文摘要 · Abstract (English)
Large language models (LLMs) increasingly support decisions about uncertain future events, yet evaluating their ability to forecast real-world outcomes remains difficult. In particular, existing benchmarks are typically static and retrospective, and therefore cannot test how information is synthesized by LLMs to predict future events under uncertainty. We introduce LLM-SoccerArena (https://llm-soccerarena.com), a prospective live benchmark that evaluates how well LLMs forecast real-world sports events before the outcomes are known. LLM-SoccerArena provides (1) a prospective live benchmark protocol, (2) a public open-source platform, and (3) a factorial benchmark design together with tournament-related questions (e.g., which team will win). LLM-SoccerArena automatically records timestamped, schema-validated forecasts of unresolved events, together with prompts, model versions, tool traces, and costs. The factorial design varies along four dimensions: (1) model version (e.g., GPT-5.5, Claude Opus 4.8); (2) information access; (3) prompting strategy, and (4) forecast horizon. We demonstrate LLM-SoccerArena through a large-scale evaluation of the 2026 FIFA World Cup, in which seven LLMs generated forecasts for all 104 matches and 15 tournament-related questions. We provide a detailed analysis of model performance across information access, prompting strategy, and forecast horizon. As a result, LLM-SoccerArena provides new evidence about the forecasting performance of state-of-the-art LLMs. For example, LLMs with web access outperform those without, but only by a small margin (i.e., a 0.023 improvement in Brier score). Overall, LLM-SoccerArena provides a flexible, open-source platform for prospective benchmarking of unresolved events. LLM-SoccerArena will be continuously updated, and can be directly applied to future national and international tournaments and league competitions.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。