用基因组嵌入预测微生物群落丰度,提升新基因组泛化能力。
Set-Aggregated Genome Embeddings for Microbiome Abundance Prediction

- 通过集合聚合基因组嵌入,利用基因组语言模型少样本学习能力。
- 在新基因组上表现优于传统生物信息学方法,提升泛化性。
- 揭示潜在表征对性能的关键作用,适合微生物组研究者参考。
微生物群落的功能编码在其全体基因组的宏基因组中。一个自然问题是:仅凭成员的原始DNA序列,能否预测群落水平的丰度特征?本文采用集合聚合基因组嵌入(SAGE)方法,利用基因组语言模型(GLMs)的少样本学习能力,预测社区级丰度谱。实验表明,该方法在新基因组上的泛化性能优于经典生物信息学方法。模型消融分析显示,群落级潜在表征直接提升了性能。最后,我们展示了潜在表征间中间变换的益处,并对比了不同GLM嵌入选择的差异。
原文摘要 · Abstract (English)
Microbiome functions are encoded within the genes of the community-wide metagenome. A natural question is whether properties of a microbial community can be predicted just from knowing the raw DNA sequences of its members. In this work, we employ set-aggregated genome embeddings (SAGE) to predict community-level abundance profiles, exploiting the few-shot learning capabilities of genomic language models (GLMs). We benchmark this approach to show improved generalization on novel genomes compared to classical bioinformatics approaches. Model ablation shows that community-level latent representations directly result in improved performance. Lastly, we demonstrate the benefits of intermediate transformations between latent representations and demonstrate the differences between GLM embedding choices.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。