arXiv:2608.03416cs.AIcs.LG2026-08

十款大模型同场预测2026世界杯,GPT-5.5 Thinking夺冠

AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction

论文配图:AI World Cup 2026: Benchmarking Large Language Models for End-to-End Football Tournament Prediction
图 1 · 摘自论文原文
  • 统一数据与规则,让10个大模型对完整世界杯赛程进行端到端预测
  • GPT-5.5 Thinking以744分领跑,唯一正确预测西班牙夺冠的模型
  • 淘汰赛表现决定最终排名,小组赛成绩几乎不影响总分

大型语言模型(LLM)常被用于预测现实事件,但因输入信息、工具使用和评估规则不同,难以比较。本文完成了首个‘AI世界杯’基准测试,十位基于LLM的助手对2026年国际足联世界杯全程进行了赛前预测。所有提交均采用相同赛事快照、提示、JSON结构和评分流程,涵盖小组积分、排名、淘汰赛对阵、最终名次、置信度及简要解释。104场比赛结束后,GPT-5.5 Thinking以744分排名第一,其次为GPT-5.5(717分)、Gemini(699分)、Qwen 3.7(687分)。仅GPT-5.5 Thinking正确预测西班牙胜阿根廷夺得冠军。最终排名主要由淘汰赛表现驱动:总分与淘汰赛得分高度相关(r=0.986),而与小组赛得分(r=0.055)、小组排名得分(r=-0.103)或两者综合(r=-0.054)关联极低。单场准确率则导致不同排序:Claude Sonnet 4.6在小组赛预测中准确率最高(63.89%),但仅列第六。平均自评置信度与结果准确性(r=-0.060)或总分(r=-0.067)无关。结果表明,完整赛事预测考验的能力不同于逐场预测,且排行榜高度依赖评分设计。评测材料、原始响应与评分代码已公开,支持复现与扩展。

原文摘要 · Abstract (English)

Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules. This paper reports the completed \emph{AI World Cup} benchmark, in which ten LLM-based assistants made a single pre-tournament forecast of the entire 2026 FIFA World Cup. Every submission used the same tournament snapshot, prompt, JSON schema, and scoring procedure. The forecasts covered group-stage scores, group rankings, the knockout bracket, final placings, confidence values, and short explanations. After all 104 matches had been played, GPT-5.5 Thinking finished first with 744 points, followed by GPT-5.5 with 717, Gemini with 699, and Qwen 3.7 with 687. GPT-5.5 Thinking was also the only model to select Spain, which defeated Argentina 1--0 in the final, as champion. The final ranking was driven mainly by knockout performance: total score was strongly correlated with knockout points ($r=0.986$), but showed little relationship with group-stage match points ($r=0.055$), group-standing points ($r=-0.103$), or their combined pre-knockout score ($r=-0.054$). Match-level accuracy produced a different ordering. Claude Sonnet 4.6 correctly predicted the largest number of group-stage outcomes (63.89\%) but placed sixth overall. Average self-reported confidence was also unrelated to either outcome accuracy ($r=-0.060$) or total score ($r=-0.067$). The results suggest that forecasting a complete tournament tests something different from predicting matches one at a time, while also showing how strongly a bracket-based leaderboard can depend on scoring design. The benchmark materials, raw responses, and scoring code are released to support replication and future extensions.

大模型评测赛事预测世界杯AI竞赛

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。