arXiv:2507.04562cs.LGcs.AI2025-07被引 4

评测大模型在真实世界预测任务中的表现,发现其仍不如专业预测者。

Evaluating LLMs on Real-World Forecasting Against Expert Forecasters

  • 在Metaculus的464个预测题上测试前沿大模型
  • 模型Brier得分看似超过普通人集体,但仍落后专家群体
  • 揭示大模型预测能力仍有明显提升空间,适合关注AI评估的研究者

大型语言模型(LLMs)在多种任务中展现出惊人能力,但其对未来事件的预测能力仍缺乏充分研究。一年前,大模型还难以接近人类群体的准确率。本文在Metaculus平台的464个预测问题上,评估了当前最先进的大模型,并与顶尖预测者进行对比。结果显示,前沿模型的Brier得分看似优于人类群体,但仍显著低于专家群体的表现。这表明尽管大模型在预测任务上取得进展,但距离专业水平仍有较大差距。

原文摘要 · Abstract (English)

Large language models (LLMs) have demonstrated remarkable capabilities across diverse tasks, but their ability to forecast future events remains understudied. A year ago, large language models struggle to come close to the accuracy of a human crowd. I evaluate state-of-the-art LLMs on 464 forecasting questions from Metaculus, comparing their performance against top forecasters. Frontier models achieve Brier scores that ostensibly surpass the human crowd but still significantly underperform a group of experts.

大模型评估预测能力专家系统

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。