构建64个细菌互作数据集,用机器学习预测微生物是竞争还是合作。
Friend or Foe
- 基于基因组尺度代谢模型生成2600万组环境下的细菌互作数据。
- 在10,000对细菌中验证了机器学习可有效预测其互作关系。
- 适合微生物生态、计算生物学及机器学习交叉研究者使用。
微生物生态学中的一个核心挑战是判断细菌在不同环境条件下是竞争还是合作。随着基因组尺度代谢模型的发展,如今我们能够以实验无法企及的规模,模拟数千对细菌在数千种环境条件下的相互作用。这些方法产生了海量数据,可被先进的机器学习算法挖掘,揭示驱动互作的机制。本文提出Friend or Foe,一个包含64个表格型环境数据集的综合性资源,涵盖超过2600万组共享环境,涉及来自两大最大代谢模型库的10,000多对细菌。该数据集专为多种机器学习任务(监督、无监督、生成)精心整理,旨在解决细菌互作背后的特定科学问题。我们对各类最新模型进行了基准测试,结果表明机器学习在此领域具有成功潜力。进一步分析还揭示了细菌互作的可预测性,并指出了新的研究方向,如细菌如何推断并导航其关系。
原文摘要 · Abstract (English)
A fundamental challenge in microbial ecology is determining whether bacteria compete or cooperate in different environmental conditions. With recent advances in genome-scale metabolic models, we are now capable of simulating interactions between thousands of pairs of bacteria in thousands of different environmental settings at a scale infeasible experimentally. These approaches can generate tremendous amounts of data that can be exploited by state-of-the-art machine learning algorithms to uncover the mechanisms driving interactions. Here, we present Friend or Foe, a compendium of 64 tabular environmental datasets, consisting of more than 26M shared environments for more than 10K pairs of bacteria sampled from two of the largest collections of metabolic models. The Friend or Foe datasets are curated for a wide range of machine learning tasks -- supervised, unsupervised, and generative -- to address specific questions underlying bacterial interactions. We benchmarked a selection of the most recent models for each of these tasks and our results indicate that machine learning can be successful in this application to microbial ecology. Going beyond, analyses of the Friend or Foe compendium can shed light on the predictability of bacterial interactions and highlight novel research directions into how bacteria infer and navigate their relationships.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。