arXiv:2502.19676cs.LGcs.CL2025-02NeurIPS被引 6

构建真实世界预测评估基准,测试模型预测与自信度。

FOReCAst: The Future Outcome Reasoning and Confidence Assessment Benchmark

  • 涵盖真假判断、时间预测、数量估计三类真实问题
  • 评估模型预测准确率与信心水平的匹配程度
  • 适合评估可信预测系统,如经济与科技趋势分析

预测在技术、经济等多个领域至关重要。然而现有预测基准普遍缺乏全面的置信度评估,问题类型有限,且多为人工构造,与真实人类预测需求脱节。为此,我们提出FOReCAst(未来结果推理与置信度评估基准),用于评估模型做出预测及其置信度的能力。该基准覆盖布尔型问题、时间范围预测和数量估计等多种现实场景,支持对预测准确性与置信度校准的综合评估,适用于真实世界应用。

原文摘要 · Abstract (English)

Forecasting is an important task in many domains, such as technology and economics. However existing forecasting benchmarks largely lack comprehensive confidence assessment, focus on limited question types, and often consist of artificial questions that do not align with real-world human forecasting needs. To address these gaps, we introduce FOReCAst (Future Outcome Reasoning and Confidence Assessment), a benchmark that evaluates models' ability to make predictions and their confidence in them. FOReCAst spans diverse forecasting scenarios involving Boolean questions, timeframe prediction, and quantity estimation, enabling a comprehensive evaluation of both prediction accuracy and confidence calibration for real-world applications.

预测评估置信度基准测试

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。