用预测区间评估大模型的量化预测能力,发现现有模型普遍高估准确性。
QuantSightBench: Evaluating LLM Quantitative Forecasting with Prediction Intervals

- 提出预测区间作为量化预测的评估框架,强调不确定性表达与校准。
- 11个前沿模型均未达90%覆盖目标,最优模型仅79.1%,且极端值处严重过自信。
- 适合关注模型可信度、量化决策风险的研究者与从业者。
预测已成为衡量大模型在不确定性下推理能力的自然基准。然而,现有评估仍局限于二元或选择题等简单判断任务。现实中,经济、公共卫生和社会人口等领域决策依赖对连续数量的数值估计,而当前基准无法捕捉此类能力。评估这类估计需显式表达不确定性并可检验。本文提出预测区间作为合适接口,要求模型具备尺度感知、多置信度一致性及连续结果上的校准能力,优于点估计。为此,我们构建新基准 QuantSightBench,评估前沿模型在多种设置下的表现,考察实际覆盖率与区间锐度。结果表明,11个被评估的前沿与开源模型均未达到90%覆盖率目标,最佳表现者 Gemini 3.1 Pro(79.1%)、Grok 4(76.4%)和 GPT-5.4(75.3%)均至少落后10个百分点。在极端量级下校准性能急剧下降,揭示所有模型普遍存在系统性过度自信。
原文摘要 · Abstract (English)
Forecasting has become a natural benchmark for reasoning under uncertainty. Yet existing evaluations of large language models remain limited to judgmental tasks in simple formats, such as binary or multiple-choice questions. In practice, however, forecasting spans a far broader scope. Across domains such as economics, public health, and social demographics, decisions hinge on numerical estimates over continuous quantities, a capability that current benchmarks do not capture. Evaluating such estimates requires a format that makes uncertainty explicit and testable. We propose prediction intervals as a natural and rigorous interface for this purpose. They demand scale awareness, internal consistency across confidence levels, and calibration over a continuum of outcomes, making them a more suitable evaluation format than point estimates for numerical forecasting. To assess this capability, we introduce a new benchmark QuantSightBench, and evaluate frontier models under multiple settings, assessing both empirical coverage and interval sharpness. Our results show that none of the 11 evaluated frontier and open-weight models achieves the 90\% coverage target, with the top performers Gemini 3.1 Pro (79.1\%), Grok 4 (76.4\%), and GPT-5.4 (75.3\%) all falling at least 10 percentage points short. Calibration degrades sharply at extreme magnitudes, revealing systematic overconfidence across all evaluated models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。