用游戏模拟世界,让AI快速预测未来状态并评估推理能力。
ForecastBench-Sim: A Simulated-World Forecasting Benchmark

- 基于文明游戏的模拟环境生成可重复的预测任务
- 支持任意时间点的连续或二元预测,且结果即时可得
- 适合研究动态世界中的概率推理,尤其适合模型调试
通用AI系统的预测基准通常受现实世界限制:结果反馈慢、极端事件罕见、反事实问题难以评分。我们提出ForecastBench-Sim,一个基于策略游戏Freeciv(类文明系列)回合制推演构建的模拟世界预测基准。预测者接收固定的世界报告(当前游戏状态的结构化快照),回答关于隐藏未来状态的问题;系统继续模拟并自动评分。由于环境为模拟,相同设定可生成任意时间跨度的连续或二元预测题,搭配干预场景以支持条件或因果问题,并可轻松生成稀有或颠覆性事件的已决样本。本文描述了基准流程、问题类型、评分机制及发布内容,并报告了模型评估验证集与匿名人类试点数据。ForecastBench-Sim旨在补充真实世界基准,提供可控、快速响应的任务,用于研究动态世界下的概率推理。
原文摘要 · Abstract (English)
Forecasting benchmarks for general-purpose AI systems usually inherit the constraints of the real world: outcomes resolve slowly, tail events are rare, and counterfactual questions are difficult to score. We introduce ForecastBench-Sim, a simulated-world forecasting benchmark built on game rollouts from Freeciv, a turn-based strategy game modelled on the Civilization series. Forecasters receive a fixed world report (a structured snapshot of the current game state) and answer questions about hidden future states; the benchmark then continues the simulation and scores forecasts. Because the world is simulated, the same setup can generate continuous or binary forecasting questions at arbitrary time horizons, paired intervention worlds for conditional or causal questions, and resolved examples of rare or disruptive outcomes. We describe the benchmark pipeline, question families, scoring protocol, and release artifacts, and report validation slices from model evaluations and an anonymized human pilot. ForecastBench-Sim is intended to complement real-world forecasting benchmarks by providing controlled, immediately resolvable tasks for studying probabilistic reasoning under dynamic world states.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。