用世界杯实时预测测试大模型,避免答案泄露。
WorldCup Arena: Prospective, Leakage-Free Evaluation of Frontier LLMs on a Live Tournament

- 在世界杯期间实时提问,无历史答案可查,杜绝数据泄露。
- 模型平均准确率63.9%,与押注热门队持平。
- 适合评估大模型真实推理能力,非记忆能力。
衡量大语言模型预测能力的基准测试几乎都是回溯性的:事件已发生,答案存在于网络,评测需防范记忆。本文提出相反设计:在2026年国际足联世界杯的39天中,六款前沿大模型(均具备扩展思维和原生服务器端网络搜索能力)在每场比赛开赛前,逐场填写104场比赛的七市场预测卡,以及12个小组冠军和赛前总冠预测;提问时答案尚不存在,因此评估天然无泄露。冻结档案共包含4,494条有评分的预测。结果显示,六种系统表现出一致行为:对比赛结果的平均准确率为63.9%,与押注庄家热门队相当;彼此共识远高于正确率,多数投票无效;低估平局与进球数,集中于单一典型比分。准确率随比赛悬殊程度变化,最激烈对决反而表现最差,而整体赛事问题回答良好。当前前沿模型在此任务上差异不大,排名稳定,中间梯队波动频繁,差距始终较小。所有背景资料、赛程及官方结果已作为基准发布,并附评分代码。
原文摘要 · Abstract (English)
Benchmarks that measure the forecasting ability of large language models are almost always retrospective: the event has happened, the answer is somewhere on the Web, and the evaluation must defend itself against memorisation. We report the opposite design. Over the 39 days of the 2026 FIFA World Cup, six frontier LLMs -- all with extended thinking and native server-side web search -- were asked before every kickoff, one match at a time, to fill in a seven-market prediction card for all 104 matches, plus 12 group winners and a pre-tournament outright pool; no answer existed when the question was asked, so the evaluation is leakage-free by construction rather than by filtering, and the frozen archive holds 4,494 scored predictions. What the tournament establishes is a set of behaviours the six systems share. On match outcome they average 63.9%, level with backing the bookmaker's favourite -- which is in fact what they usually do. They agree with one another far more often than they are right, so a majority vote adds nothing. They under-commit to draws and to goals, and crowd their scoreline picks onto a single prototypical result. Accuracy tracks how lopsided a fixture is rather than how much is known about it: it collapses in the closest ties, where the dossiers are richest, while questions about the tournament as a whole are answered well. On this task the current generation of frontier systems is not sharply differentiated: the standings hold up at the top and the bottom across the run and churn in the middle, and the margins stay narrow throughout. The briefing dossiers, fixtures and official results are released as a benchmark, together with the scoring code.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。