基于稀疏贝叶斯的多任务模型,可从微生物组数据中精准识别关键菌群并量化预测不确定性。
Hierarchical Sparse Bayesian Multitask Model with Scalable Inference for Microbiome Analysis
- 构建分层稀疏贝叶斯多任务模型,共享跨任务的稀疏结构以提升泛化能力。
- 在合成数据上实现更优的支持恢复,在真实微生物组数据中准确识别特征菌群。
- 适用于异质数据融合,适合需可信预测与不确定性分析的生物医学研究者。
本文提出一种分层贝叶斯多任务学习模型,适用于具有共享稀疏结构的多任务二分类问题。基于变分推断设计高效推理算法,近似后验分布。在多种合成数据集和基于微生物组谱型预测人类健康状态的任务上验证了该方法的有效性。分析整合了来自多个微生物组研究的数据,并与多种基准方法进行了全面比较。合成数据实验表明,当不同任务间回归系数存在共同稀疏结构时,所提方法具备更优的支持恢复性能。真实微生物组分类实验显示,该方法能有效提取有信息量的分类单元,提供校准良好的预测结果与不确定性量化,且在预测指标上表现竞争力。值得注意的是,尽管数据来源存在异质性(如实验目标、实验室设置、测序设备、人群特征差异),该方法仍表现出稳健性能。
原文摘要 · Abstract (English)
This paper proposes a hierarchical Bayesian multitask learning model that is applicable to the general multi-task binary classification learning problem where the model assumes a shared sparsity structure across different tasks. We derive a computationally efficient inference algorithm based on variational inference to approximate the posterior distribution. We demonstrate the potential of the new approach on various synthetic datasets and for predicting human health status based on microbiome profile. Our analysis incorporates data pooled from multiple microbiome studies, along with a comprehensive comparison with other benchmark methods. Results in synthetic datasets show that the proposed approach has superior support recovery property when the underlying regression coefficients share a common sparsity structure across different tasks. Our experiments on microbiome classification demonstrate the utility of the method in extracting informative taxa while providing well-calibrated predictions with uncertainty quantification and achieving competitive performance in terms of prediction metrics. Notably, despite the heterogeneity of the pooled datasets (e.g., different experimental objectives, laboratory setups, sequencing equipment, patient demographics), our method delivers robust results.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。