用可验证奖励训练模型预测现实事件,准确率超前沿模型且更靠谱。
Outcome-based Reinforcement Learning to Predict the Future
- 基于可验证奖励的强化学习,让模型从真实事件结果中学习。
- 140亿参数模型在预测任务上超越o1,校准性显著提升。
- 适合需要高可信度预测的金融、政策等场景使用。
强化学习结合可验证奖励(RLVR)已在编程与数学等领域有效提升大语言模型的推理能力。本文将其应用于预测真实世界事件——这一任务因结果噪声大、延迟高而极具挑战。基于一个来自预测市场的近期问题数据集及关联新闻标题,我们训练了一个140亿参数的推理模型,其预测准确率可媲美甚至超过前沿模型o1,同时大幅改善概率校准效果。在Polymarket交易模拟中,该模型的投注在测试集上预估可实现超10%的投资回报率。我们详细比较了训练方法,包括用合成预测问题扩充数据、引入学习稳定性约束,以及推理时采用中位数预测采样策略。
原文摘要 · Abstract (English)
Reinforcement Learning with Verifiable Rewards (RLVR) has been an effective approach for improving Large Language Models' reasoning in domains such as coding and mathematics. Here, we apply RLVR methods towards forecasting future real-world events - a challenging task for RL due to the very noisy (and delayed) outcomes involved. Using a novel dataset of recent questions from a prediction market, and accompanying relevant news headlines, we show that a compact (14B) reasoning model can be trained to match or surpass the predictive accuracy of frontier models like o1, while greatly improving probabilistic calibration. The model's performance is also practically meaningful: in a Polymarket trading simulation, we estimate that its bets would have yielded a return on investment of over 10% across all questions in the test set. We detail and compare approaches used in training our model, including augmenting our training-data with synthetic prediction questions, guardrails for learning stability, and median prediction sampling at inference-time.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。