用真实观测数据评估7个AI天气模型,发现它们在季风期表现不错但仍有明显误差。
MAUSAM: An Observations-focused assessment of Global AI Weather Prediction Models During the South Asian Monsoon
- 基于地面站、雨量计和卫星影像,直接对比观测数据评估模型性能。
- 模型对极端降水和气旋路径等小尺度现象仍存在系统性低估,误差比再分析高15%-45%。
- 适合关注区域天气预测、模型评估与数据稀疏区应用的研究者阅读。
精准天气预报对社会规划和防灾至关重要,但在观测稀疏地区仍具挑战。当前人工智能(AI)天气预测评估主要依赖再分析数据,可能掩盖关键缺陷。本文提出MAUSAM(南亚季风期间测量AI不确定性),利用地面气象站、雨量网络和静止卫星影像,对七个主流AI天气模型——FourCastNet、FourCastNet-SFNO、Pangu-Weather、GraphCast、Aurora、AIFS和GenCast——在南亚季风期的表现进行观测驱动评估。结果显示,这些模型在大尺度温压风及降水、云量和次季节至季节尺度涡旋统计等方面表现出色,体现数据驱动预测的优势;然而在细尺度上仍存在系统误差,如极端降水低估、气旋路径偏差和中尺度动能谱失真。与再分析相比,模型相对于观测的误差高出15%-45%,表明以再分析为基准会高估模型性能。其中AIFS在大气变量表征上最一致,GraphCast与GenCast也展现强预测能力。该研究构建了区域预测评估框架,揭示了AI天气预测在数据稀疏区的潜力与局限,强调观测评估对未来业务化应用的重要性。
原文摘要 · Abstract (English)
Accurate weather forecasts are critical for societal planning and disaster preparedness. Yet these forecasts remain challenging to produce and evaluate, especially in regions with sparse observational coverage. Current evaluation of artificial intelligence (AI) weather prediction relies primarily on reanalyses, which can obscure important deficiencies. Here we present MAUSAM (Measuring AI Uncertainty during South Asian Monsoon), an evaluation of seven leading AI-based forecasting systems - FourCastNet, FourCastNet-SFNO, Pangu-Weather, GraphCast, Aurora, AIFS, and GenCast - during the South Asian Monsoon, using ground-based weather stations, rain gauge networks, and geostationary satellite imagery. The AI models demonstrate impressive forecast skill during monsoon across a broad range of variables, ranging from large-scale surface temperature and winds to precipitation, cloud cover, and subseasonal to seasonal eddy statistics, highlighting the strength of data-driven weather prediction. However, the models still exhibit systematic errors at finer scales like the underprediction of extreme precipitation, divergent cyclone tracks, and the mesoscale kinetic energy spectra, highlighting avenues for future improvement. A comparison against observations reveals forecast errors up to 15-45% larger than those relative to reanalysis and traditional forecasts, indicating that reanalysis-centric benchmarks can overstate forecast skill. Of the models assessed, AIFS achieves the most consistent representation of atmospheric variables, with GraphCast and GenCast also showing strong skill. The analysis presents a framework for evaluating AI weather models on regional prediction and highlights both the promise and current limitations of AI weather prediction in data-sparse regions, underscoring the importance of observational evaluation for future operational adoption.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。