arXiv:2511.18394cs.LG2025-11被引 2

大模型预测能力因问题类型和提问方式差异巨大

Future Is Unevenly Distributed: Forecasting Ability of LLMs Depends on What We're Asking

  • 测试不同模型在真实事件预测中的表现,分析提问方式影响
  • 模型准确率随问题领域和上下文变化显著,最高差达30%
  • 适合关注模型预测边界与提示工程的研究者

大型语言模型(LLMs)在社会、政治和经济事件中展现出部分预测能力,但其预测表现随领域结构和提示框架的差异而剧烈波动。我们研究了不同模型家族在模型训练截止日期后真实事件预测任务中的表现,分析了上下文、问题类型和外部知识对准确率与校准性的影响,并探讨了添加新闻事实如何改变模型信念形成与失效模式。结果表明,预测能力高度依赖于所问内容及提问方式。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate partial forecasting competence across social, political, and economic events. Yet, their predictive ability varies sharply with domain structure and prompt framing. We investigate how forecasting performance varies with different model families on real-world questions about events that happened beyond the model cutoff date. We analyze how context, question type, and external knowledge affect accuracy and calibration, and how adding factual news context modifies belief formation and failure modes. Our results show that forecasting ability is highly variable as it depends on what, and how, we ask.

大模型预测提示工程认知偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。