提出新型微生物群落计数模型,解决稀疏数据下的物种推断难题。
Poisson Hierarchical Indian Buffet Processes-With Indications for Microbiome Species Sampling Models
- 基于泊松层次印度茶点过程,实现跨组与组内信息共享。
- 可处理无限物种数量,精准区分技术性与生物性零值。
- 适合微生物组、基因组及文本分析中的层次计数建模。
我们提出泊松层次印度茶点过程(PHIBP),一种新型物种抽样模型,用于应对复杂稀疏计数数据的挑战,通过在组间与组内实现信息共享来提升建模能力。理论发展构建了可进行贝叶斯非参数推断的框架,具备机器学习特性,能够从数据中学习潜在无限数量的物种(分类单元)参数。聚焦于微生物组分析,该模型填补了关键空白:提供灵活的多变量计数模型,考虑过度分散,并稳健处理多种数据类型(如OTUs、ASVs)。引入反映物种丰度与多样性的新参数。模型在组间借力同时明确区分技术性与生物性零值,以解释稀疏共现模式。结果是具备可计算后验推断、精确生成采样及对未见物种问题的严谨解决方案。我们还描述了扩展形式,允许领域专家通过协变量和结构化先验融入知识,支持菌株级分析。虽起源于生态学,本工作为遗传学、商业与文本分析中的层次计数建模提供了通用方法,对概率统计中物种抽样理论具有深远影响。
原文摘要 · Abstract (English)
We introduce the Poisson Hierarchical Indian Buffet Process (PHIBP), a new class of species sampling models designed to address the challenges of complex, sparse count data by facilitating information sharing across and within groups. Our theoretical developments enable a tractable Bayesian nonparametric framework with machine learning elements, accommodating a potentially infinite number of species (taxa) whose parameters are learned from data. Focusing on microbiome analysis, we address key gaps by providing a flexible multivariate count model that accounts for overdispersion and robustly handles diverse data types (OTUs, ASVs). We introduce novel parameters reflecting species abundance and diversity. The model borrows strength across groups while explicitly distinguishing between technical and biological zeros to interpret sparse co-occurrence patterns. This results in a framework with tractable posterior inference, exact generative sampling, and a principled solution to the unseen species problem. We describe extensions where domain experts can incorporate knowledge through covariates and structured priors, with potential for strain-level analysis. While motivated by ecology, our work provides a broadly applicable methodology for hierarchical count modeling in genetics, commerce, and text analysis, and has significant implications for the broader theory of species sampling models arising in probability and statistics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。