arXiv:2501.16008stat.MEcs.LG2025-01被引 1

提出新方法估算未知物种数,计算更快且更准。

Gaussian credible intervals in Bayesian nonparametric estimation of the unseen

  • 基于贝叶斯非参数模型,用高斯近似推导可信区间
  • 无需蒙特卡洛采样,计算效率更高,误差更小
  • 适合大规模生物数据中的物种多样性分析

未知物种问题假设从包含不同物种(可能无限)的种群中采集了 n≥1 个样本,要求估计若再采集 m≥1 个新样本时会发现的未见物种数 K_{n,m}。本文在 Pitman-Yor 先验下采用贝叶斯非参数方法,提出一种针对大 m 值的 K_{n,m} 渐近可信区间构造新方法。通过利用后验分布的高斯中心极限定理,该方法可完整参数化 Pitman-Yor 先验(包括 Dirichlet 先验),且无需蒙特卡洛采样,显著提升计算效率。在合成与真实数据上的验证表明,该方法能显著缩小渐近与精确可信区间之间的差距,对任意 m≥1 均有更好表现。

原文摘要 · Abstract (English)

The unseen-species problem assumes $n\geq1$ samples from a population of individuals belonging to different species, possibly infinite, and calls for estimating the number $K_{n,m}$ of hitherto unseen species that would be observed if $m\geq1$ new samples were collected from the same population. This is a long-standing problem in statistics, which has gained renewed relevance in biological and physical sciences, particularly in settings with large values of $n$ and $m$. In this paper, we adopt a Bayesian nonparametric approach to the unseen-species problem under the Pitman-Yor prior, and propose a novel methodology to derive large $m$ asymptotic credible intervals for $K_{n,m}$, for any $n\geq1$. By leveraging a Gaussian central limit theorem for the posterior distribution of $K_{n,m}$, our method improves upon competitors in two key aspects: firstly, it enables the full parameterization of the Pitman-Yor prior, including the Dirichlet prior; secondly, it avoids the need of Monte Carlo sampling, enhancing computational efficiency. We validate the proposed method on synthetic and real data, demonstrating that it improves the empirical performance of competitors by significantly narrowing the gap between asymptotic and exact credible intervals for any $m\geq1$.

贝叶斯统计物种多样性可信区间非参数

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。