整合多源数据,用深度学习预测加州70多种作物的产量。
California Crop Yield Benchmark: Combining Satellite Image, Climate, Evapotranspiration, and Soil Data Layers for County-Level Yield Forecasting of Over 70 Crops
- 融合卫星、气候、蒸散和土壤数据,构建多模态模型。
- 在未见数据上实现0.76的R2,跨区域预测性能强。
- 适合农业气象、精准农业与气候适应研究者使用。
加利福尼亚州是全球农业生产的领军地区,占美国总产量的12.5%,在全球食品与棉花供应中位列第五。尽管美国农业部国家农业统计局提供了丰富的历史产量数据,但受环境、气候与土壤因素复杂交互影响,准确及时的作物产量预测仍具挑战。本研究构建了一个覆盖全加州所有县区、涵盖70余种作物、时间跨度为2008至2022年的综合性作物产量基准数据集。该数据集融合了Landsat卫星影像、日尺度气候记录、月度蒸散数据及高分辨率土壤属性。为有效处理异构输入,我们开发了一种面向县级、作物特定的多模态深度学习模型,采用分层特征提取与时间序列编码器捕捉生长季内的时空动态;静态输入如土壤特性与作物种类用于建模长期变异。模型在未见测试数据上整体达到0.76的R²得分,展现出在加州多样化农业区域中的强大预测能力。该基准与建模框架为农业预测、气候适应与精准农业的发展提供了重要基础。完整数据集与代码库已公开于我们的GitHub仓库。
原文摘要 · Abstract (English)
California is a global leader in agricultural production, contributing 12.5% of the United States total output and ranking as the fifth-largest food and cotton supplier in the world. Despite the availability of extensive historical yield data from the USDA National Agricultural Statistics Service, accurate and timely crop yield forecasting remains a challenge due to the complex interplay of environmental, climatic, and soil-related factors. In this study, we introduce a comprehensive crop yield benchmark dataset covering over 70 crops across all California counties from 2008 to 2022. The benchmark integrates diverse data sources, including Landsat satellite imagery, daily climate records, monthly evapotranspiration, and high-resolution soil properties. To effectively learn from these heterogeneous inputs, we develop a multi-modal deep learning model tailored for county-level, crop-specific yield forecasting. The model employs stratified feature extraction and a timeseries encoder to capture spatial and temporal dynamics during the growing season. Static inputs such as soil characteristics and crop identity inform long-term variability. Our approach achieves an overall R2 score of 0.76 across all crops of unseen test dataset, highlighting strong predictive performance across California diverse agricultural regions. This benchmark and modeling framework offer a valuable foundation for advancing agricultural forecasting, climate adaptation, and precision farming. The full dataset and codebase are publicly available at our GitHub repository.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。