融合检索增强与多尺度建模,提升大区域玉米产量预测精度
Retrieval-Augmented Multi-scale Framework for County-Level Crop Yield Prediction Across Large Regions
- 构建多尺度模型捕捉日级生长动态与跨年长期依赖
- 在630个县数据上实现比基线更优的预测表现
- 适合需要高鲁棒性农业预测的政策制定者和保险机构
本文提出一种新型作物产量预测方法,对制定管理策略、评估保险风险及保障长期粮食安全至关重要。尽管现有数据驱动方法在此领域已展现潜力,但在大范围地理区域和长时间跨度下性能常下降。这主要源于两大挑战:(1)难以同时捕捉短期与长期时间模式;(2)无法有效应对农业系统中的空间数据异质性。忽略这些问题会导致特定区域或年份预测不可靠,影响政策与资源分配。为此,本文提出新框架:首先设计一种新骨干模型,同时建模日级作物生长动态与跨年长期依赖;为提升跨区域泛化能力,引入基于检索的适应策略。针对年度间产量差异,设计新颖的检索-精炼流程,通过剔除输入特征无法解释的跨年偏差来优化检索样本。在覆盖美国630个县的真实县级玉米产量数据上的实验表明,该方法持续优于多种基线模型,验证了检索增强策略在空间异质性下的鲁棒性提升效果。
原文摘要 · Abstract (English)
This paper proposes a new method for crop yield prediction, which is essential for developing management strategies, informing insurance assessments, and ensuring long-term food security. Although existing data-driven approaches have shown promise in this domain, their performance often degrades when applied across large geographic regions and long time periods. This limitation arises from two key challenges: (1) difficulty in jointly capturing short-term and long-term temporal patterns, and (2) inability to effectively accommodate spatial data variability in agricultural systems. Ignoring these issues often leads to unreliable predictions for specific regions or years, which ultimately affects policy decisions and resource allocation. In this paper, we propose a new predictive framework to address these challenges. First, we introduce a new backbone model architecture that captures both short-term daily-scale crop growth dynamics and long-term dependencies across years. To further improve generalization across diverse spatial regions, we augment this model with a retrieval-based adaptation strategy. Recognizing the substantial yield variation across years, we design a novel retrieval-and-refinement pipeline that adjusts retrieved samples by removing cross-year bias not explained by input features. Our experiments on real-world county-level corn yield data over 630 counties in the US demonstrate that our method consistently outperforms different types of baselines. The results also verify the effectiveness of the retrieval-based augmentation method in improving model robustness under spatial heterogeneity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。