用已知结果的过去事件测试大模型预测能力,实现可重复的预测评估。
Bench to the Future: A Pastcasting Benchmark for Forecasting Agents
- 通过已知答案的‘回溯预测’任务,模拟真实预测场景。
- 在数十万网页数据支持下,大模型表现与实时预测接近。
- 适合研究模型预测能力演进或评估推理方法的研究者。
预测是一项具有明确衡量标准的挑战性任务,可用于研究人工智能系统。由于预测需大量网络调研,且评估依赖事件实际发生时间,构建预测基准极为困难。目前尚无基准能为大语言模型(LLM)提供真实、封闭且可重复的预测环境。本文提出Bench To the Future(BTF),一个‘回溯预测’基准,包含数百个高质量问题,其结果已知。每个问题均配有数万条相关网页组成的离线语料库,使大模型能在历史事件上生成逼真的‘预测’。实验表明,该环境产生的结果与基于互联网实时未决问题的预测相当。我们使用多个LLM(包括新发布的Claude 4模型)评测了代理式和思维链预测方法,并展示了BTF追踪预测能力随时间演进的能力。该基准将作为动态更新的活体基准,持续新增问题以应对训练数据截止日期的推移。欢迎研究人员通过[email protected]联系获取使用支持。
原文摘要 · Abstract (English)
Forecasting is a challenging task that offers a clearly measurable way to study AI systems. Forecasting requires a large amount of research on the internet, and evaluations require time for events to happen, making the development of forecasting benchmarks challenging. To date, no forecasting benchmark provides a realistic, hermetic, and repeatable environment for LLM forecasters. We introduce Bench To the Future (BTF), a "pastcasting" benchmark with hundreds of high-quality questions for which the resolution is already known. Each question is accompanied by a large offline corpus of tens of thousands of relevant web pages, enabling a way to elicit realistic "forecasts" on past events from LLMs. Results suggest that our pastcasting environment can produce results comparable to those based on forecasts using the internet on at-the-time unresolved questions. We show results benchmarking agent and chain-of-thought forecasting approaches using several LLMs, including the recently-released Claude 4 models, and demonstrate BTF's ability to track steady forecasting capability progress over time. We intend this to be a living benchmark, with new questions added continually to account for increasing training data cutoff dates. We invite researchers to contact us at [email protected] to utilize our benchmark or tooling for their own research.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。