arXiv:2605.22672cs.AI2026-05

更强大的语言模型在复杂预测中反而表现更差,尤其在高风险场景下。

Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most

论文配图:Is Capability a Liability? More Capable Language Models Make Worse Forecasts When It Matters Most
图 1 · 摘自论文原文
  • 模型越强,在超线性增长和突发变化任务中越偏离真实分布
  • 能力越强的模型过度上移尾部预测,低估实际风险
  • 现有评估指标会掩盖问题,应改用连续无界评分

我们在具有超线性增长和制度突变尾部风险的时间序列预测任务中发现大语言模型存在反向缩放现象,这类结构常见于金融与流行病学。在这些任务中,模型能力越强,其分布预测越差。该现象在我们发布的无污染模拟基准 ForecastBench-Sim(FBSim)上得到验证,包括匹配线性控制的合成SIR流行病预测,且在真实数据集如新冠疫情、麻疹、房地产市场和超通胀中均可复现。分位数分解显示失败集中于上尾,能力强的模型为追踪激进外推而过度抬高上尾,下尾则基本不变。对 Llama-3.1 的族内研究发现,模型规模与后训练均独立贡献此效应。领域知识无法可靠改善校准。该反向缩放在单阈值指标中不显现,反而使能力-准确率关系符号反转;常规阈值评分会忽略上尾代价,而包含尾部的评分可逆转同一输出的能力-准确率关系。我们建议在大语言模型预测评估中同时使用连续(无界)准确性度量与有界二值阈值指标。

原文摘要 · Abstract (English)

We document inverse scaling in LLMs on forecasting problems whose underlying time series exhibit superlinear growth and tail risk of regime change, a structure common in finance and epidemiology. On these tasks, more capable models produce worse distributional forecasts. The pattern appears on ForecastBench-Sim (FBSim), a contamination-free, simulated-world benchmark we release, in forecasting synthetic SIR epidemics with a matched linear control, and replicates in real-world datasets on COVID-19, measles, housing markets, and hyperinflation. A per-quantile decomposition shows the failure concentrates at the upper tail, which more capable models shift upward to track aggressive extrapolations of growth, while the lower tail stays put. A within-family study of Llama-3.1 shows that both model scale and post-training independently contribute to this effect. Domain knowledge does not reliably rescue calibration. This inverse scaling does not appear on single-threshold metrics common in LLM forecasting benchmarks, reversing the sign of the capability--accuracy relationship on identical outputs. Single-threshold scoring at conventional cutoffs misses the upper-tail cost; tail-inclusive scoring reverses the sign of the capability--accuracy relationship on the same outputs. We recommend that LLM forecasting evaluations use continuous (and unbounded) measures of accuracy alongside bounded binary threshold metrics.

大模型预测反向缩放风险建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。