arXiv:2502.08622cs.LG2025-02被引 2

用机器学习预测加州干旱,LSTM模型表现最优。

Forecasting Drought Using Machine Learning in California

  • 采用LSTM、XGBoost等模型预测干旱等级,以历史数据为输入。
  • 模型在12周预报中MAE仅0.33,分类F1达0.9,精度高。
  • 适合干旱模式稳定、严重干旱频发的地区使用。

干旱是加州频繁且代价高昂的自然灾害,对农业生产和水资源(尤其是地下水)造成重大影响。本研究评估了多种机器学习方法在预测加州美国干旱监测(USDM)分类中的表现,包括卷积神经网络(CNN)、随机森林、XGBoost和长短期记忆(LSTM)循环神经网络,并与基准持续性模型进行对比。通过宏观F1二分类指标评估模型对严重干旱(USDM等级D2或更高)的预测能力。结果显示,LSTM模型表现最佳,其次为XGBoost、CNN和随机森林。县级层面分析表明,LSTM在干旱模式一致且严重干旱较常见的地区表现更优,而在干旱评分快速上升区域表现较差。利用30周历史数据,LSTM成功预测未来12周干旱程度,平均绝对误差(MAE)为0.33,相当于小于半级的误差(0-5级量表)。此外,其宏观F1得分为0.9,表明对严重干旱的二分类准确率高。不同时间窗口与未来预测时长的评估显示,至少需24周数据以达最佳性能,且短时预测(尤其少于8周)效果更佳。

原文摘要 · Abstract (English)

Drought is a frequent and costly natural disaster in California, with major negative impacts on agricultural production and water resource availability, particularly groundwater. This study investigated the performance of applying different machine learning approaches to predicting the U.S. Drought Monitor classification in California. Four approaches were used: a convolutional neural network (CNN), random forest, XGBoost, and long short term memory (LSTM) recurrent neural network, and compared to a baseline persistence model. We evaluated the models' performance in predicting severe drought (USDM drought category D2 or higher) using a macro F1 binary classification metric. The LSTM model emerged as the top performer, followed by XGBoost, CNN, and random forest. Further evaluation of our results at the county level suggested that the LSTM model would perform best in counties with more consistent drought patterns and where severe drought was more common, and the LSTM model would perform worse where drought scores increased rapidly. Utilizing 30 weeks of historical data, the LSTM model successfully forecasted drought scores for a 12-week period with a Mean Absolute Error (MAE) of 0.33, equivalent to less than half a drought category on a scale of 0 to 5. Additionally, the LSTM achieved a macro F1 score of 0.9, indicating high accuracy in binary classification for severe drought conditions. Evaluation of different window and future horizon sizes in weeks suggested that at least 24 weeks of data would result in the best performance, with best performance for shorter horizon sizes, particularly less than eight weeks.

干旱预测LSTM机器学习加州

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。