arXiv:2412.19217cs.LGstat.ML2024-12被引 2

用神经网络自动学特征,让物种分布模型更准

Applying the maximum entropy principle to neural networks enhances multi-species distribution models

  • 用神经网络替代人工设计特征,自动提取共享模式
  • 在六个区域六类生物上均优于Maxent等主流模型
  • 特别适合采样不均的复杂场景,提升预测精度

公民科学推动了生物多样性数据库快速增长,尤其是仅有存在记录(PO)的数据。这类数据对物种分布建模至关重要,但受限于采样偏差和缺乏缺失信息。当前广泛使用泊松点过程构建物种分布模型(SDM),其中最大熵方法(Maxent)通过最大化概率分布熵来建模。然而,传统方法依赖人工设计特征。本文提出DeepMaxent,利用神经网络自动从复杂输入中学习跨物种共享特征,基于归一化泊松损失函数,对每个物种的出现概率由神经网络建模。在存在空间采样偏差的基准数据集上,使用PO数据校准、PA数据验证,覆盖六个不同生物类群与环境变量区域。结果表明,DeepMaxent在所有区域和类群中均优于Maxent及其他领先模型,尤其在采样不均地区表现突出,显著提升预测准确性,优于传统单物种模型,为方法改进提供新可能。

原文摘要 · Abstract (English)

The rapid expansion of citizen science initiatives has led to a significant growth of biodiversity databases, and particularly presence-only (PO) observations. PO data are invaluable for understanding species distributions and their dynamics, but their use in a Species Distribution Model (SDM) is curtailed by sampling biases and the lack of information on absences. Poisson point processes are widely used for SDMs, with Maxent being one of the most popular methods. Maxent maximises the entropy of a probability distribution across sites as a function of predefined transformations of variables, called features. In contrast, neural networks and deep learning have emerged as a promising technique for automatic feature extraction from complex input variables. Arbitrarily complex transformations of input variables can be learned from the data efficiently through backpropagation and stochastic gradient descent (SGD). In this paper, we propose DeepMaxent, which harnesses neural networks to automatically learn shared features among species, using the maximum entropy principle. To do so, it employs a normalised Poisson loss where for each species, presence probabilities across sites are modelled by a neural network. We evaluate DeepMaxent on a benchmark dataset known for its spatial sampling biases, using PO data for calibration and presence-absence (PA) data for validation across six regions with different biological groups and covariates. Our results indicate that DeepMaxent performs better than Maxent and other leading SDMs across all regions and taxonomic groups. The method performs particularly well in regions of uneven sampling, demonstrating substantial potential to increase SDM performances. In particular, our approach yields more accurate predictions than traditional single-species models, which opens up new possibilities for methodological enhancement.

物种分布深度学习最大熵生物多样性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。