arXiv:2602.11771cs.AI2026-02

提出新方法提升多物种分布模型的二值化精度,避免误判稀有物种。

How to Optimize Multispecies Set Predictions in Presence-Absence Modeling ?

  • 基于评估指标直接优化,选择最可能的物种组合作为预测结果
  • 在强类别不平衡和高稀有度下表现优于传统阈值法
  • 适合生态学、保护规划等需要准确物种共现信息的研究

物种分布模型(SDMs)通常生成概率性出现预测,需转化为二值存在-缺失图以用于生态推断与保护规划。然而,此二值化步骤常为启发式处理,会显著扭曲物种普遍性与群落组成估计。本文提出 MaxExp,一种决策驱动的二值化框架,通过直接最大化选定评估指标来选取最可能的物种组合。MaxExp 无需校准数据,适用于多种评分标准。我们还引入计算高效的 Set Size Expectation(SSE)方法,基于预期物种丰富度预测群落。在涵盖不同类群、物种数量和性能指标的三个案例研究中,MaxExp 均稳定优于或媲美广泛使用的阈值法与校准方法,尤其在强类别不平衡和高稀有度情况下表现突出。SSE 提供更简单的替代方案,且具备竞争力。二者共同提供稳健、可复现的多物种 SDM 二值化工具。

原文摘要 · Abstract (English)

Species distribution models (SDMs) commonly produce probabilistic occurrence predictions that must be converted into binary presence-absence maps for ecological inference and conservation planning. However, this binarization step is typically heuristic and can substantially distort estimates of species prevalence and community composition. We present MaxExp, a decision-driven binarization framework that selects the most probable species assemblage by directly maximizing a chosen evaluation metric. MaxExp requires no calibration data and is flexible across several scores. We also introduce the Set Size Expectation (SSE) method, a computationally efficient alternative that predicts assemblages based on expected species richness. Using three case studies spanning diverse taxa, species counts, and performance metrics, we show that MaxExp consistently matches or surpasses widely used thresholding and calibration methods, especially under strong class imbalance and high rarity. SSE offers a simpler yet competitive option. Together, these methods provide robust, reproducible tools for multispecies SDM binarization.

物种分布二值化生态建模多物种

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。