arXiv:2608.28482cs.LGcs.AI2026-08

不同评分规则影响大模型预测的误差结构,选择关键。

How Proper Scoring Rules Shape LLM Forecasting

论文配图:How Proper Scoring Rules Shape LLM Forecasting
图 1 · 摘自论文原文
  • 用五种正确评分规则训练大模型做二元预测
  • 布里尔规则模型表现最好,对数规则校准最优
  • 相同准确率下,错误类型差异显著,适合调参研究

本文评估了奖励函数选择如何影响大模型预测者的表现与行为。比较了五种正确评分规则作为二元预测任务的训练目标,针对真实世界事件的已解决预测。尽管这些规则在理论上均激励诚实的概率报告,但训练出的模型在校准度、概率使用和偏差、信息、噪声的估计分布上存在差异,总体准确率和区分能力差异较小。布里尔规则训练的模型观测到最低的布里尔得分和最高的AUC-ROC,对数规则训练的模型则拥有最高的对数得分和最低的校准误差。表现相近的模型通过不同的偏差、信息和噪声组合达成性能,表明正确评分规则在训练目标上并非可互换。奖励选择不仅影响大模型预测效果,也塑造其预测误差的结构。每种条件仅使用单一随机种子,部分差异可能源于训练随机性。

原文摘要 · Abstract (English)

This paper evaluates how reward function choice shapes the performance and behavior of LLM forecasters. We compare five proper scoring rules as training objectives for binary forecasts of resolved real-world events. Although the rules share the same theoretical incentive for truthful probability reporting, the resulting models differ in calibration, probability use, and estimated profiles of bias, information, and noise, with smaller differences in aggregate accuracy and discrimination. The Brier-trained model has the lowest observed Brier score and highest AUC-ROC, while the log-trained model has the highest observed log score and lowest calibration error. Models with similar aggregate performance also reach that performance through different combinations of bias, information, and noise. Proper scoring rules therefore need not behave interchangeably as training objectives. Reward choice may shape not only how well an LLM forecasts, but how its forecasting errors are structured. Each condition uses a single seed, so some differences may reflect training stochasticity.

大模型预测评分规则误差分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。