arXiv:2601.22444cs.LGcs.AI2026-01被引 2

用大模型自动生成并验证真实世界预测题,提升AI评估效率与质量。

Automating Forecasting Question Generation and Resolution for AI Evaluation

  • 基于大模型网络调研代理,自动生成高多样性真实预测问题。
  • 问题可验证率96%,解决准确率达95%,优于人工平台Metaculus。
  • 可用于评估预测模型性能,优化策略如分解法显著降低误差。

预测未来事件在决策中极具价值,是衡量通用智能的有力指标。由于预测具有概率性,开发和评估AI预测模型需大量多样且困难的问题,并准确判断结果。以往自动化方法依赖重复数据源(如天气、股票),限制了问题多样性与实用性。本文提出一个基于大模型网络研究代理的系统,可大规模自动生成并验证高质量预测问题。我们生成了1499个多样化的真实世界预测问题,并在数月后进行验证。系统生成可验证、无歧义问题的比例约为96%,超过领先的人工标注平台Metaculus。问题解决准确率约为95%。实验表明,使用更智能的大模型驱动的预测代理表现更优:Gemini 3 Pro Brier得分为0.134,GPT-5为0.149,Gemini 2.5 Flash为0.179。此外,我们展示了该系统可直接用于改进预测——在生成的问题集上评估问题分解策略,使Brier得分从0.141提升至0.132。

原文摘要 · Abstract (English)

Forecasting future events is highly valuable in decision-making and is a robust measure of general intelligence. As forecasting is probabilistic, developing and evaluating AI forecasters requires generating large numbers of diverse and difficult questions, and accurately resolving them. Previous efforts to automate this laborious work relied on recurring data sources (e.g., weather, stocks), limiting diversity and utility. In this work, we present a system for generating and resolving high-quality forecasting questions automatically and at scale using LLM-powered web research agents. We use this system to generate 1499 diverse, real-world forecasting questions, and to resolve them several months later. We estimate that our system produces verifiable, unambiguous questions approximately 96% of the time, exceeding the rate of Metaculus, a leading human-curated forecasting platform. We also find that our system resolves questions at approximately 95% accuracy. We verify that forecasting agents powered by more intelligent LLMs perform better on these questions (Brier score of 0.134 for Gemini 3 Pro, 0.149 for GPT-5, and 0.179 for Gemini 2.5 Flash). Finally, we demonstrate how our system can be leveraged to directly improve forecasting, by evaluating a question decomposition strategy on a generated question set, yielding a significant improvement in Brier scores (0.132 vs. 0.141).

AI评估预测生成大模型应用自动化评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。