arXiv:2509.23936cs.CL2025-09被引 1

测试大模型能否随新信息更新预测,发现其反应保守且不一致。

Do Language Models Update their Forecasts with New Information?

  • 设计EvolveCast框架,评估模型对训练后新信息的响应能力。
  • 模型更新预测时普遍保守,信心估计与人类差距大。
  • 适合关注大模型推理可信度与知识更新机制的研究者。

以往研究多将预测视为静态任务,未考虑新证据出现时预测及其置信度应如何演化。为此,我们提出EvolveCast框架,用于评估大语言模型是否能适当根据训练截止日期后的信息更新其预测。通过人类预测者作为参照,评估模型在新信息下的预测修正与置信度校准表现。结果显示,尽管模型对新信息有一定响应,但更新行为常不一致或过于保守。无论是口头表述还是基于逻辑值的置信度估计,均远未达到人类基准水平。在多种大模型和设置下,模型普遍表现出预测更新的保守倾向。这表明当前基于RAG等方法的知识更新策略不足以支持概率推理;模型将新信息视为检索上下文而非改变后验概率的证据。EvolveCast强调了建立更稳健机制以融入外部知识并动态调整信念的必要性。

原文摘要 · Abstract (English)

Prior work has largely treated forecasting as a static task, failing to consider how forecasts and the confidence in them should evolve as new evidence emerges. To address this gap, we introduce EvolveCast, a framework for evaluating whether large language models revise their forecasts appropriately in response to new information. In particular, EvolveCast assesses whether LLMs update their forecasts when presented with information released after their training cutoff. We use human forecasters as a comparative reference to assess forecast updates and confidence calibration under new information. While LLMs demonstrate some responsiveness to new information, their updates are often inconsistent or overly conservative. We further find that both verbalized and logits-based confidence estimates remain far from the human reference standard. Across settings with a variety of LLMs, models tend to be conservative in updating their forecasts. These findings suggest that current approaches (e.g., RAG-based methods) for updating model knowledge are insufficient for probabilistic reasoning; models treat new information as retrieval context rather than evidence that shifts posterior probability. EvolveCast thus underscores the need for more robust mechanisms to incorporate external knowledge into belief dynamics.

大模型推理预测更新置信度校准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。