用强化学习让大模型突破知识截止,精准预测真实事件。
Reinforcement Learning for LLM-based Event Forecasting

- 用GRPO强化学习微调1.5B-14B参数模型,接入实时信息源。
- 1.5B模型在交叉熵上超越Claude Sonnet 3.5,性能更优。
- 揭示了大模型预测的可扩展性与不确定性边界,适合决策研究者。
我们采用最近提出的高效样本与内存的强化学习方法Group Relative Policy Optimization (GRPO),对1.5B至14B参数的预训练大语言模型进行微调。这些模型通过维基百科修订工具或新闻摘要获取实时信息,以预测超出其知识截止时间的真实事件,以及模拟不同动态特性的任务。基于实验结果,我们探讨了大模型在预测任务中的可扩展性,并根据固有随机不确定性(如掷骰子)将判断型预测归类至可验证/不可验证范畴。经GRPO训练后,1.5B参数的Qwen 2.5模型在相同数据集上的交叉熵表现优于Claude Sonnet 3.5。同时,我们也讨论了实现该成果过程中的多个失败路径。
原文摘要 · Abstract (English)
We use Group Relative Policy Optimization (GRPO), a recently devised sample and memory efficient reinforcement learning method, to finetune pretrained LLMs in the range of 1.5B to 14B parameters equipped with the ability to get current information through the use of a Wikipedia revisions tool, or news summaries, to forecast real events beyond the knowledge cutoff of the LLM, as well as problems made to simulate different aspects of the dynamics of that training. We use the results of these experiments to comment on the scaling capability of LLMs for forecasting, as well as classify how judgmental forecasting fits into the verifiable/unverifiable domain taxonomy, considering the impact of the inherent aleatoric uncertainty when forecasting future events (e.g. the roll of a die). As a result of the GRPO training, we manage to bring a 1.5B parameter transformer (Qwen 2.5 1.5B) to forecasting performance superior to Claude Sonnet 3.5 over the same dataset as measured by cross entropy from the market agreed probabilities. We also discuss various dead ends on the path to this result.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。