arXiv:2503.13057cs.LGcs.CV2025-03被引 7

MaskSDM通过掩码训练提升物种分布模型的灵活、鲁棒与可解释性。

MaskSDM with Shapley values to improve flexibility, robustness, and explainability in species distribution modeling

  • 采用掩码训练策略,支持推理时任意选择输入变量而不需重训
  • 在12,738种植物上表现优于插补方法,接近特定子集模型性能
  • 结合Shapley值精准量化变量贡献,适合生态研究与保护规划

物种分布模型(SDMs)在生物多样性研究、保护规划和生态位建模中至关重要,通过环境条件预测物种分布。预测因子的选择对模型准确性和生态模式反映至关重要。然而现有模型,包括传统与深度学习方法,普遍缺乏三个关键能力:(i) 推理时灵活选择相关预测因子而无需重训;(ii) 对缺失预测值具备鲁棒性;(iii) 可解释性以准确量化各预测因子贡献。为此,我们提出MaskSDM,一种基于深度学习的新型SDM,通过掩码训练实现灵活的变量选择,可在任意输入子集下进行预测,并对缺失数据保持鲁棒。同时,该模型利用Shapley值精确评估各变量贡献,优于传统近似方法。我们在全球sPlotOpen数据集上对12,738种植物进行建模,结果表明MaskSDM性能超越插补方法,且逼近仅用特定子集训练的模型。这些发现凸显其在提升SDMs适用性与推广潜力方面的价值,为构建可广泛应用于生态研究的领域基础模型奠定基础。

原文摘要 · Abstract (English)

Species Distribution Models (SDMs) play a vital role in biodiversity research, conservation planning, and ecological niche modeling by predicting species distributions based on environmental conditions. The selection of predictors is crucial, strongly impacting both model accuracy and how well the predictions reflect ecological patterns. To ensure meaningful insights, input variables must be carefully chosen to match the study objectives and the ecological requirements of the target species. However, existing SDMs, including both traditional and deep learning-based approaches, often lack key capabilities for variable selection: (i) flexibility to choose relevant predictors at inference without retraining; (ii) robustness to handle missing predictor values without compromising accuracy; and (iii) explainability to interpret and accurately quantify each predictor's contribution. To overcome these limitations, we introduce MaskSDM, a novel deep learning-based SDM that enables flexible predictor selection by employing a masked training strategy. This approach allows the model to make predictions with arbitrary subsets of input variables while remaining robust to missing data. It also provides a clearer understanding of how adding or removing a given predictor affects model performance and predictions. Additionally, MaskSDM leverages Shapley values for precise predictor contribution assessments, improving upon traditional approximations. We evaluate MaskSDM on the global sPlotOpen dataset, modeling the distributions of 12,738 plant species. Our results show that MaskSDM outperforms imputation-based methods and approximates models trained on specific subsets of variables. These findings underscore MaskSDM's potential to increase the applicability and adoption of SDMs, laying the groundwork for developing foundation models in SDMs that can be readily applied to diverse ecological applications.

物种分布可解释性深度学习生态建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。