arXiv:2608.09972physics.ao-phcs.AI2026-08

AI天气模型并非普遍在极端天气上表现差,部分模型甚至更优。

Do AI weather models miss extremes?

  • 对比11种物理与AI模型,评估其在欧洲多地的极端天气预测能力。
  • 多个AI模型在强风、高温、晴天太阳辐射等极端条件下领先,最高提升24.8%。
  • 模型缺陷是特定的,非AI整体问题,数值预报模型也有类似偏差。

首批人工智能天气模型常被指在极端天气下表现不佳,主要源于基于再分析数据的确定性回归系统评估。本文将十一类物理与AI预报系统,在欧洲气象站(包括地表、太阳辐射和雨量站)连续十个月的数据上进行验证,评估10米风速、2米气温、小时短波辐射累积及小时降水,以均方误差(MAE)对比ECMWF IFS在ERA5 1991–2020气候区间的表现。结果显示,AI模型在极端尾部并无统一相对技能劣势:Jua EPT-2.1 Europa在全条件风速上领先8.4%,Jua EPT-2 HRRR在温度整体及高温条件下分别领先12.1%和19.6%±2.2%;两者及DWD ICON Global在烈风条件下最优。Jua EPT-2.1 Helios在太阳辐射总体及云层遮蔽和晴空尾部分别领先10.2%±1.7%、16.4%±3.4%和24.8%±5.4%。降水方面,三个Jua模型在中等强度时提升14–15%,在P75–P95区间提升9–11%;而EPT-2 Reasoning在超95百分位仍领先1.7%±0.5%。失败情况为模型特异:ECMWF AIFS在高温尾部下降4.9%±2.0%,NOAA GFS下降22.8%±2.0%。所有模型(含数值预报)均表现出向观测分布中心的共性偏差,模型间差异远小于共同信号。因此,极端天气相对性能缺失并非AI模型共性,而是特定模型问题。

原文摘要 · Abstract (English)

First-generation AI weather models are often reported to underperform at extremes, mostly in reanalysis-based evaluations of deterministic regression systems. We verify eleven physical and AI forecast systems against European synoptic, solar, and rain-gauge stations over ten months for 10 m wind, 2 m temperature, hourly shortwave accumulation, and hourly precipitation, scoring mean absolute error (MAE) against ECMWF IFS in ERA5 1991-2020 climatological regimes. Among these systems, AI models do not show a uniform relative-skill deficit in the tails. Jua EPT-2.1 Europa leads all-conditions wind (+8.4%), while Jua EPT-2 HRRR leads temperature overall (+12.1%) and in the heat regime (+19.6 +/- 2.2%). EPT-2.1 Europa and DWD ICON Global lead at gale-force wind. Jua EPT-2.1 Helios leads solar overall (+10.2 +/- 1.7%), in overcast conditions (+16.4 +/- 3.4%), and in the clear-sky tail (+24.8 +/- 5.4%). For precipitation, three Jua models gain 14-15% at moderate intensity and 9-11% at P75-P95; EPT-2 Reasoning remains ahead above P95 (+1.7 +/- 0.5%). Failures are model-specific: ECMWF AIFS loses 4.9 +/- 2.0% in the heat tail, while NOAA GFS loses 22.8 +/- 2.0% there. Every model, including numerical weather prediction systems, shows a shared conditional bias toward the centre of the observed distribution, with an inter-model spread several times smaller than the shared signal. Missing relative skill at extremes is therefore not a property of AI weather models as a class, but of particular AI and physical models.

AI天气极端天气模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。