arXiv:2601.06336cs.LGcs.AI2026-01被引 9

用真实世界结果做监督,让大模型自动学预测。

Future-as-Label: Scalable Supervision from Real-World Outcomes

  • 用可验证事件结果当奖励,训练语言模型做概率预测。
  • 在真实预测任务上Brier分数提升27%,校准误差减半。
  • 无需人工标注,适合开放世界长期预测场景。

时间带来免费的监督信号:对现实事件的预测最终会得到可验证的结果。我们拓展强化学习框架,将可验证的奖励机制应用于长期现实预测。训练语言模型基于因果掩码信息做出概率预测,使用合适的评分规则作为奖励函数,一旦事件结果确定即更新模型。整个学习过程完全依赖实际发生的结果,实现大规模、基于结果的开放世界监督。在真实世界预测基准测试中,采用预见学习(Foresight Learning)训练的Qwen3-32B相比其预训练基线,Brier得分提升27%,校准误差减半;尽管参数量仅为Qwen3-235B的1/7,仍优于后者在构造未来事件预测任务和Metaculus基准上的表现。

原文摘要 · Abstract (English)

Time creates free supervision: forecasts about real-world events resolve to verifiable outcomes. The passage of time provides labels that require no annotation. To exploit this structure, we extend reinforcement learning with verifiable rewards to real-world prediction over time. We train language models to make probabilistic forecasts from causally masked information, using proper scoring rules as the reward function once events resolve. Learning is driven entirely by realized outcomes, enabling scalable outcome-based supervision in open-world prediction. On real-world forecasting benchmarks, Qwen3-32B trained using Foresight Learning improves Brier score by 27% and halves calibration error relative to its pretrained baseline, and outperforms Qwen3-235B on both constructed future-event prediction tasks and the Metaculus benchmark despite a 7x parameter disadvantage.

预测模型强化学习自监督语言模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。