构建动态预测基准,评估AI在未知未来事件上的预测能力。
ForecastBench: A Dynamic Benchmark of AI Forecasting Capabilities
- 用自动生成的1000个未来事件问题构成动态评测集,确保无数据泄露。
- 在200个随机问题上测试,专家预测显著优于顶尖大模型(p<0.001)。
- 适合关注AI预测能力、可解释性及人类优势的研究者使用。
对未来事件的预测是决策的重要依据。机器学习系统有潜力大规模提供预测,但目前缺乏对ML系统在标准化预测问题上准确性的评估框架。为此,我们提出ForecastBench:一个动态基准,通过自动创建并定期更新的1000个预测问题,评估机器学习系统的准确性。所有问题均针对尚未发生且无已知答案的未来事件,杜绝数据泄露可能。我们通过收集专家、公众及大语言模型(LLMs)在基准中随机抽取的200个问题上的预测,量化当前系统的表现。尽管大模型在多数基准上表现超人,但在本任务中,专家预测显著优于最先进大模型(p值<0.001)。系统与人类得分将公开展示于www.forecastbench.org。
原文摘要 · Abstract (English)
Forecasts of future events are essential inputs into informed decision-making. Machine learning (ML) systems have the potential to deliver forecasts at scale, but there is no framework for evaluating the accuracy of ML systems on a standardized set of forecasting questions. To address this gap, we introduce ForecastBench: a dynamic benchmark that evaluates the accuracy of ML systems on an automatically generated and regularly updated set of 1,000 forecasting questions. To avoid any possibility of data leakage, ForecastBench is comprised solely of questions about future events that have no known answer at the time of submission. We quantify the capabilities of current ML systems by collecting forecasts from expert (human) forecasters, the general public, and LLMs on a random subset of questions from the benchmark ($N=200$). While LLMs have achieved super-human performance on many benchmarks, they perform less well here: expert forecasters outperform the top-performing LLM ($p$-value $<0.001$). We display system and human scores in a public leaderboard at www.forecastbench.org.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。