arXiv:2604.26106cs.AI2026-04

通过可复现的推理追踪,揭示预测模型的策略性思维差异。

Evaluating Strategic Reasoning in Forecasting Agents

论文配图:Evaluating Strategic Reasoning in Forecasting Agents
图 1 · 摘自论文原文
  • 构建1417个回溯问题数据集,支持可复现的研究与预测
  • 检测出0.004的贝叶斯分数差异,区分研究与判断能力
  • 发现优秀预测者更擅长预判盲点和黑天鹅事件

预测基准测试仅提供准确率排行榜,难以揭示为何某些预测者更优。我们引入Bench to the Future 2(BTF-2),包含1,417个回溯问题及一个冻结的1500万文档研究语料库,使智能体可在离线环境下可复现地进行研究与预测,并生成完整的推理轨迹。BTF-2可检测出0.004的贝叶斯分数差异,能区分智能体在研究与判断上的不同优势。我们构建的预测器比任何单一前沿模型高出0.011的贝叶斯分数,并用于无事后偏见地评估智能体的战略推理能力。结果表明,表现更优的预测者主要在于其对自身盲点的预审分析以及对黑天鹅事件的考量。专家人类预测者指出,前沿智能体的主要战略推理失误集中在评估政治与商业领袖的动机、判断其执行既定计划的可能性,以及建模制度运作过程方面。

原文摘要 · Abstract (English)

Forecasting benchmarks produce accuracy leaderboards but little insight into why some forecasters are more accurate than others. We introduce Bench to the Future 2 (BTF-2), 1,417 pastcasting questions with a frozen 15M-document research corpus in which agents reproducibly research and forecast offline, producing full reasoning traces. BTF-2 detects accuracy differences of 0.004 Brier score, and can distinguish differential agent strengths in research vs. judgment. We build a forecaster 0.011 Brier more accurate than any single frontier agent, and use it to evaluate agent strategic reasoning without hindsight bias. We find the better forecaster differs primarily in its pre-mortem analysis of its blind spots and consideration of black swans. Expert human forecasters found the dominant strategic reasoning failures of frontier agents are in assessing political and business leaders' incentives, judging their likelihood to follow through on stated plans, and modeling institutional processes.

预测模型战略推理贝叶斯评分黑天鹅

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。