arXiv:2608.17971cs.LG2026-08

对比机器学习与深度学习模型在极端干旱下的玉米产量预测表现

Evaluating and improving crop-yield forecasting methods during extreme drought

论文配图:Evaluating and improving crop-yield forecasting methods during extreme drought
图 1 · 摘自论文原文
  • 用16个气象变量做预测,比较非深度与深度学习模型
  • 极端干旱年份数据超出历史范围,导致预测误差增大
  • 样本加权和特征选择提升非深度学习模型效果

气候变化对粮食生产的影响催生了多种基于机器学习(ML)、数值天气预报(NWP)或混合式ML-NWP的预测模型,用于揭示气象驱动因素与作物生长之间的结构与物理关系,以预测作物产量。例如2012年美国中西部玉米带干旱即为极端事件,严重冲击作物生产并考验预测模型极限。本研究使用16个气象变量作为预测因子,对比非深度学习与深度学习模型在2012年极端干旱年份的县级别玉米产量预测表现。该问题特征在于训练集与测试集的特征分布存在显著差异——极端干旱年的气象条件超出了历史观测范围。此外,数据存在时空不规则性:部分县缺失产量数据导致空间稀疏,每年仅使用部分日度数据导致时间稀疏。为此,采用样本加权与特征选择进行模型改进。结果显示,这些改进显著提升了非深度学习模型性能,但深度学习模型VITA改善有限。尽管VITA整体优于非深度学习模型,本研究揭示了训练与测试特征分布不一致对预测模型的影响,对比了深度学习与非深度学习模型表现,并验证了适用于非深度学习模型的有效改进策略。

原文摘要 · Abstract (English)

The impact of climate variability on food production has led to the creation of various forecasting models that uses machine learning (ML), numerical weather predictors (NWP) or a hybrid of ML-NWP models to identify structural and physical relationships between meteorological drivers and crop growth, in order to predict crop yield. Droughts, for example the 2012 Midwestern US (Corn Belt) drought, are extreme events that affect crop production and test the limits of these forecasting models. Using 16 meteorological drivers as predictors, we compare ML (non-deep learning) and deep learning forecasting models to predict the county-level corn yield for the extreme drought year, 2012. This forecasting problem is characterized by a dissimilarity between the feature distributions of the training and test data, where the meteorological conditions of the extreme drought year fall outside the range of historically observed values. Additionally, the dataset consists of spatial and temporal irregularities where counties with missing yields introduce spatial sparsity and the use of only a subset of daily values per year introduce temporal sparsity. To overcome this, we use sample weighting and feature selection as modifications to improve our forecasting models. These modifications lead to an improvement for ML models; however, the deep learning model VITA shows little to no improvement. While VITA outperforms the ML models with or without modifications, our current study sheds light on the effect of dissimilarity between train and test feature distributions on forecasting models, compares deep learning versus non-deep learning models, and introduces modifications that are effective for non-deep learning models.

产量预测干旱影响机器学习模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。