针对物种分布模型的交叉验证偏差问题,提出基于空间自相关的稳健评估方法。
Foundation for unbiased cross-validation of spatio-temporal models for species distribution modeling
- 采用空间、时空阻隔的交叉验证设计,避免因邻近观测相关性导致结果偏高
- 随机交叉验证使AUC虚高最多达0.16,且均方误差高出80%以上
- 推荐使用基于实际空间自相关范围的分块策略与外部时间验证
评估物种分布模型(SDMs)在真实部署场景下的预测性能,需妥善处理数据中的空间与时间依赖。交叉验证(CV)是标准评估方法,但其设计显著影响结果有效性。当模型用于空间或时间迁移时,随机交叉验证因相邻观测的空间自相关(SAC)会导致结果过度乐观。本文在温带植物和溯河鱼类两个真实存在-缺失数据集上,对比了四种机器学习算法(GBM、XGBoost、LightGBM、Random Forest),采用多种交叉验证设计:随机、空间、时空、环境及前向链式。评估两种训练策略(LAST FOLD 和 RETRAIN),并在每种方案内进行超参数调优。模型性能通过独立的出时测试集以AUC、MAE和相关系数衡量。结果显示,随机交叉验证使AUC虚高最多达0.16,且均方误差最高比空间分块方案高出80%。在经验空间自相关范围进行分块可显著降低偏差。训练策略影响评估结果:强空间自相关下,LAST FOLD的验证-测试差异更小;弱空间自相关时,RETRAIN获得更高测试AUC。集成提升模型在空间结构化交叉验证中表现最优。建议采用基于空间自相关感知的分块、分块内调参与外部时间验证的稳健工作流,以提升模型在空间与时间迁移下的可靠性。
原文摘要 · Abstract (English)
Evaluating the predictive performance of species distribution models (SDMs) under realistic deployment scenarios requires careful handling of spatial and temporal dependencies in the data. Cross-validation (CV) is the standard approach for model evaluation, but its design strongly influences the validity of performance estimates. When SDMs are intended for spatial or temporal transfer, random CV can lead to overoptimistic results due to spatial autocorrelation (SAC) among neighboring observations. We benchmark four machine learning algorithms (GBM, XGBoost, LightGBM, Random Forest) on two real-world presence-absence datasets, a temperate plant and an anadromous fish, using multiple CV designs: random, spatial, spatio-temporal, environmental, and forward-chaining. Two training data usage strategies (LAST FOLD and RETRAIN) are evaluated, with hyperparameter tuning performed within each CV scheme. Model performance is assessed on independent out-of-time test sets using AUC, MAE, and correlation metrics. Random CV overestimates AUC by up to 0.16 and produces MAE values up to 80 percent higher than spatially blocked alternatives. Blocking at the empirical SAC range substantially reduces this bias. Training strategy affects evaluation outcomes: LAST FOLD yields smaller validation-test discrepancies under strong SAC, while RETRAIN achieves higher test AUC when SAC is weaker. Boosted ensemble models consistently perform best under spatially structured CV designs. We recommend a robust SDM workflow based on SAC-aware blocking, blocked hyperparameter tuning, and external temporal validation to improve reliability under spatial and temporal shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。