用机器学习优化气象预报,提升降水、温度和风速预测精度。
Improvements to the post-processing of weather forecasts using machine learning and feature selection

- 基于相关性分析筛选特征,用LightGBM改进预报模型。
- 在多个地点和预报时效下,模型RMSE低于原始预报和官方产品。
- 针对降水数据特点,采用特殊损失函数提升高降雨量预测表现。
本研究利用日本气象厅(JMA)提供的18个不同地理区域(平原、山区、岛屿)的中尺度模型(MSM)数据集,开发并改进基于机器学习的降水量、温度和风速后处理模型。通过引入目标位置周边网格点的气象变量作为输入特征,并基于相关性分析进行特征选择,实验表明,在当前设置下,基于LightGBM的模型相比测试的神经网络基线(包括复现的CNN模型),以及原始的MSM预报和日本气象厅后处理产品MSM Guidance(MSMG),在多数地点和预报时效上均取得了更低的均方根误差(RMSE)。由于降水量分布高度偏斜且零值较多,还额外考察了基于Tweedie损失函数和事件加权训练策略的改进方法,显著提升了高阈值降雨事件的预测性能,尽管整体表现仍略低于MSMG,且效果具有站点依赖性。
原文摘要 · Abstract (English)
This study aims to develop and improve machine learning-based post-processing models for precipitation, temperature, and wind speed predictions using the Mesoscale Model (MSM) dataset provided by the Japan Meteorological Agency (JMA) for 18 locations across Japan, including plains, mountainous regions, and islands. By incorporating meteorological variables from grid points surrounding the target locations as input features and applying feature selection based on correlation analysis, we found that, in our experimental setting, the LightGBM-based models achieved lower RMSE than the specific neural-network baselines tested in this study, including a reproduced CNN baseline, and also generally achieved lower RMSE than both the raw MSM forecasts and the JMA post-processing product, MSM Guidance (MSMG), across many locations and forecast lead times. Because precipitation has a highly skewed distribution with many zero cases, we additionally examined Tweedie-based loss functions and event-weighted training strategies for precipitation forecasting. These improved event-oriented performance relative to the original LightGBM model, especially at higher rainfall thresholds, although the gains were site dependent and overall performance remained slightly below MSMG.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。