arXiv:2512.25070cs.LGcs.CL2025-12被引 17

用新闻数据训练大模型预测未来,效果媲美大型闭源模型。

Scaling Open-Ended Reasoning to Predict the Future

  • 从每日新闻自动生成开放性预测问题,构建大规模训练集
  • 模型在2025年5-8月测试中准确率与校准度显著提升
  • 开源全部资源,推动语言模型预测研究普及

高风险决策需要对未来不确定性进行推理。本文训练语言模型回答开放性预测问题。为扩大训练数据,我们利用每日新闻中的全球事件,通过全自动、精细化的筛选流程生成新型预测问题,并基于此构建数据集OpenForesight。训练使用Qwen3思维模型,为防止训练和评估时泄露未来信息,系统全程依赖离线新闻语料库进行数据生成与检索。在小规模验证集引导下,我们验证了检索机制的有效性及强化学习中改进奖励函数的优势。最终模型OpenForecaster 8B在2025年5月至8月的留出测试中,表现与更大规模的专有模型相当,且预测准确性、校准度和一致性均得到提升。我们发现,经过预测训练带来的校准改进可泛化至多个主流基准测试。所有模型、代码与数据均开源,以促进语言模型预测研究的广泛可及性。

原文摘要 · Abstract (English)

High-stakes decision making involves reasoning under uncertainty about the future. In this work, we train language models to make predictions on open-ended forecasting questions. To scale up training data, we synthesize novel forecasting questions from global events reported in daily news, using a fully automated, careful curation recipe. We train the Qwen3 thinking models on our dataset, OpenForesight. To prevent leakage of future information during training and evaluation, we use an offline news corpus, both for data generation and retrieval in our forecasting system. Guided by a small validation set, we show the benefits of retrieval, and an improved reward function for reinforcement learning (RL). Once we obtain our final forecasting system, we perform held-out testing between May to August 2025. Our specialized model, OpenForecaster 8B, matches much larger proprietary models, with our training improving the accuracy, calibration, and consistency of predictions. We find calibration improvements from forecasting training generalize across popular benchmarks. We open-source all our models, code, and data to make research on language model forecasting broadly accessible.

未来预测语言模型自动化数据开源

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。