arXiv:2502.05253cs.CLcs.AI2025-02被引 8

让大模型通过自我对弈提升未来预测能力,无需人工标注。

LLMs Can Teach Themselves to Better Predict the Future

  • 用模型自生成推理路径和预测结果进行自我对抗训练
  • 使Phi-4 14B和DeepSeek-R1 14B预测准确率提升7%-10%
  • 适合想提升模型推理与预测能力的研究者

我们提出一种以结果为导向的微调框架,无需依赖人工标注的推理样本,即可增强大语言模型(LLM)的预测能力。方法利用模型自对弈生成多样化的推理轨迹与概率预测,针对在模型知识截止日期后才可验证的问题。随后通过实际结果距离对推理对进行排序,并采用直接偏好优化(DPO)微调模型。在独立测试集上,该方法使Phi-4 14B和DeepSeek-R1 14B的预测准确率比基础模型和使用随机标签微调的控制模型提升7%–10%,达到与更大规模前沿模型如GPT-4o相当的预测水平。

原文摘要 · Abstract (English)

We present an outcome-driven fine-tuning framework that enhances the forecasting capabilities of large language models (LLMs) without relying on human-curated reasoning samples. Our method leverages model self-play to generate pairs of diverse reasoning trajectories and probabilistic forecasts for a set of diverse questions that resolve after the models' knowledge cutoff date. We then rank pairs of these reasoning traces by their distance to the actual outcomes before fine-tuning the model via Direct Preference Optimization (DPO). On a separate test set, our approach increases prediction accuracy of Phi-4 14B and DeepSeek-R1 14B by between 7--10\% over a base model and a DPO fine-tuned control model with randomized labels, bringing them on par with forecasting capabilities of much larger frontier models like GPT-4o.

预测模型自学习大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。