arXiv:2510.17638cs.AIcs.CL2025-10被引 32

用大模型预测未来事件,发现其既有潜力也有明显短板。

LLM-as-a-Prophet: Understanding Predictive Intelligence with Prophet Arena

  • 构建实时预测评估平台,拆解预测流程进行系统测试。
  • 模型预测误差小、信心稳定,部分表现接近市场水平。
  • 适合关注大模型预测能力与局限的研究者和从业者。

预测不仅是基础性认知追求,对金融、经济等社会系统也至关重要。随着在互联网规模数据上训练的大语言模型(LLMs)快速发展,利用其预测真实世界未来事件成为可能,我们称之为「大模型即先知」(LLM-as-a-Prophet)。本文系统研究了此类预测智能。为此,我们构建了Prophet Arena——一个持续收集实时预测任务并将其分解为不同阶段的通用评估基准,支持受控且大规模实验。全面评估表明,许多大模型已具备出色的预测能力,例如较小的校准误差、一致的预测置信度以及具有前景的市场回报。然而我们也发现关键瓶颈:如模型事件召回不准确、对数据来源理解错误,以及在临近结果揭晓时信息整合速度慢于市场。

原文摘要 · Abstract (English)

Forecasting is not only a fundamental intellectual pursuit but also is of significant importance to societal systems such as finance and economics. With the rapid advances of large language models (LLMs) trained on Internet-scale data, it raises the promise of employing LLMs to forecast real-world future events, an emerging paradigm we call "LLM-as-a-Prophet". This paper systematically investigates such predictive intelligence of LLMs. To this end, we build Prophet Arena, a general evaluation benchmark that continuously collects live forecasting tasks and decomposes each task into distinct pipeline stages, in order to support our controlled and large-scale experimentation. Our comprehensive evaluation reveals that many LLMs already exhibit impressive forecasting capabilities, reflected in, e.g., their small calibration errors, consistent prediction confidence and promising market returns. However, we also uncover key bottlenecks towards achieving superior predictive intelligence via LLM-as-a-Prophet, such as LLMs' inaccurate event recalls, misunderstanding of data sources and slower information aggregation compared to markets when resolution nears.

大模型预测智能评估未来预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。