arXiv:2512.15140cs.LG2025-12被引 1

检验机器学习模型在德国作物产量预测中的泛化能力与解释可靠性。

Generalization and Feature Attribution in Machine Learning Models for Crop Yield and Anomaly Prediction in Germany

  • 对比树模型与深度学习模型在时空分割下的表现差异。
  • 模型在时间独立验证中性能显著下降,暴露泛化缺陷。
  • 强调解释结果需经严格泛化验证,尤其适用于农业数据科学。

本研究考察了用于预测德国NUTS-3区域作物产量及异常的机器学习模型的泛化性能与可解释性。基于高质量长期数据集,系统比较了集成树模型(XGBoost、随机森林)与深度学习方法(LSTM、TCN)在常规空间划分测试集与时间独立验证年份上的表现。尽管所有模型在空间划分测试中表现良好,但在时间上独立的验证年份中性能显著下降,揭示出持续存在的泛化局限。值得注意的是,某些测试集准确率高但时间验证表现差的模型仍能产生看似可信的SHAP特征重要性值,暴露出事后可解释性方法的关键脆弱性:即使模型无法泛化,解释结果也可能显得可靠。这强调了在农业与环境系统中进行验证感知的解释必要性。特征重要性不应未经验证就直接采纳,除非模型明确证明可在未见时空条件下泛化。研究倡导领域感知的验证策略、混合建模方法,并对可解释性方法提出更严格的审视。最终回应环境数据科学中的核心挑战:如何以足够鲁棒的方式评估泛化能力,以信任模型解释?

原文摘要 · Abstract (English)

This study examines the generalization performance and interpretability of machine learning (ML) models used for predicting crop yield and yield anomalies in Germany's NUTS-3 regions. Using a high-quality, long-term dataset, the study systematically compares the evaluation and temporal validation behavior of ensemble tree-based models (XGBoost, Random Forest) and deep learning approaches (LSTM, TCN). While all models perform well on spatially split, conventional test sets, their performance degrades substantially on temporally independent validation years, revealing persistent limitations in generalization. Notably, models with strong test-set accuracy, but weak temporal validation performance can still produce seemingly credible SHAP feature importance values. This exposes a critical vulnerability in post hoc explainability methods: interpretability may appear reliable even when the underlying model fails to generalize. These findings underscore the need for validation-aware interpretation of ML predictions in agricultural and environmental systems. Feature importance should not be accepted at face value unless models are explicitly shown to generalize to unseen temporal and spatial conditions. The study advocates for domain-aware validation, hybrid modeling strategies, and more rigorous scrutiny of explainability methods in data-driven agriculture. Ultimately, this work addresses a growing challenge in environmental data science: how can we evaluate generalization robustly enough to trust model explanations?

作物产量模型泛化可解释性农业AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。