用大模型分析新闻预测事件,性能媲美人类顶尖预测者。
AIA Forecaster: Technical Report
- 通过智能搜索新闻+多代理协调+统计校准提升预测准确性
- 在ForecastBench上达到人类超级预测者水平,超越以往大模型基线
- 适合对可解释预测系统、人机协作决策感兴趣的团队
本技术报告介绍AIA Forecaster——一个基于大语言模型的判断性预测系统,利用非结构化数据进行预测。该方法融合三大核心组件:对高质量新闻源的代理式搜索、协调同一事件不同预测结果的监督代理,以及用于纠正大模型行为偏差的统计校准技术。在ForecastBench基准(Karger et al., 2024)上,AIA Forecaster表现等同于人类超级预测者,优于先前的大模型基线。此外,我们引入了一个源自流动性预测市场的更具挑战性的新基准。尽管在该基准上其表现不及市场共识,但将AIA Forecaster与市场共识结合的集成模型优于市场共识本身,证明其提供了互补信息。本工作建立了人工智能预测的新基准,并为未来研究提供可复用的实用建议。据我们所知,这是首个可验证地实现规模化专家级预测的成果。
原文摘要 · Abstract (English)
This technical report describes the AIA Forecaster, a Large Language Model (LLM)-based system for judgmental forecasting using unstructured data. The AIA Forecaster approach combines three core elements: agentic search over high-quality news sources, a supervisor agent that reconciles disparate forecasts for the same event, and a set of statistical calibration techniques to counter behavioral biases in large language models. On the ForecastBench benchmark (Karger et al., 2024), the AIA Forecaster achieves performance equal to human superforecasters, surpassing prior LLM baselines. In addition to reporting on ForecastBench, we also introduce a more challenging forecasting benchmark sourced from liquid prediction markets. While the AIA Forecaster underperforms market consensus on this benchmark, an ensemble combining AIA Forecaster with market consensus outperforms consensus alone, demonstrating that our forecaster provides additive information. Our work establishes a new state of the art in AI forecasting and provides practical, transferable recommendations for future research. To the best of our knowledge, this is the first work that verifiably achieves expert-level forecasting at scale.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。