arXiv:2607.17765cs.LGcs.AI2026-07被引 1

用2026世界杯104场真实比赛,测试大模型预测能力与决策水平。

FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches

论文配图:FIFA World Cup 2026 as a Contamination-Free Benchmark for LLM Forecasting Agents: Four Models, a Bookmaker, and 104 Matches
图 1 · 摘自论文原文
  • 四款顶级大模型在无污染环境中自主完成搜索-决策-复盘循环
  • 模型预测准确率相近,但投资回报率差异大,最高达+10%最低-18%
  • 揭示模型自知之明、市场参考程度等决策质量维度的深层差异

我们提出WC2026-Agents,一个用于评估大语言模型作为自主预测代理在真实未来事件上的基准与数据集。针对2026年世界杯104场比赛,四种前沿模型——Claude Opus 4.8、ChatGPT(GPT-5.5,高推理)、Gemini 3.1 Pro和Grok(专家模式)——执行相同的搜索-行动-反思循环:使用网络工具收集证据,提交1X2(主胜/平局/客胜)概率分布及虚拟100美元投注,赛后仅基于最终比分进行反思。由于每场比赛均在模型训练截止后举行,该基准天然无污染。关键的是,我们将四个代理与来自同一信息环境的第五个竞争者——赛前博彩市场——进行对比,获取逐场比赛的1X2赔率,提供经济上合理的基线,使我们不仅能评估预测结果,还能衡量其金钱表现。数据集包含416次预测和414次反思,附完整推理文本、真实结果(含点球大战)、赔率及可复现的评估套件。基准评估显示:原始准确率掩盖了真相——四模型在92%比赛中选择相同首选;无一超越市场Brier得分;且对市场热门的固定投注收益高于所有四模型。然而,它们在决策层面显著分化:投资回报率介于-18%至+10%,所有模型在跟随市场时均无法盈利,引用市场信息的比例从12%到100%不等,错误选择的自我报告错误率在36%至86%之间。该基准因此衡量了校准度、决策质量与自我认知——这些维度即便在预测一致时也存在显著差异。

原文摘要 · Abstract (English)

We introduce WC2026-Agents, a benchmark and dataset for evaluating large language models (LLMs) as autonomous forecasting agents on real, future events. For every one of the 104 matches of the 2026 FIFA World Cup, four frontier models -- Claude Opus 4.8, ChatGPT (GPT-5.5, high reasoning), Gemini 3.1 Pro, and Grok (Expert Mode) -- ran an identical search-act-reflect loop: gather evidence with a web tool, commit to a 1X2 (team-A win / draw / team-B win) distribution and a virtual 100-USD bet, and, after the match, reflect given only the final score. Because every match kicked off after the models' training cutoffs, the benchmark is contamination-free by construction. Crucially, we pair the four agents with a fifth competitor drawn from the same information environment -- the pre-match betting market -- collected as per-match 1X2 odds, giving an economically grounded baseline and letting us score not just what an agent predicts but what it does with money. The release contains 416 forecasts and 414 reflections with verbatim reasoning, ground truth (including penalty shootouts), odds, and a reproducible evaluation suite. A reference evaluation surfaces findings that raw accuracy hides: the four agents issue an identical top pick in 92% of matches and none beats the market's Brier score; indeed, a naive flat stake on the market favorite out-earns all four agents. Yet the agents diverge sharply as decision-makers: betting return-on-investment ranges from -18% to +10%, fading the market is unprofitable for all four, the share of forecasts that cite the market ranges from 12% to 100%, and self-reported error rates on wrong picks range from 36% to 86%. The benchmark thus measures calibration, decision quality, and self-knowledge -- axes on which frontier models differ even when their predictions do not. Data and code: https://github.com/graphuofm/FIFA2026LLM

大模型评估决策分析赛事预测智能代理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。