用基因组大模型直接从原始序列生成微生物组表示,无需注释
MetagenBERT: a Transformer-based Architecture using Foundational genomic Large Language Models for novel Metagenome Representation
- 基于DNABERT2和DNABERTMS嵌入序列,用聚类生成元数据
- 在5个肠道菌群数据集上表现优于或接近物种丰度基线
- 仅用10%序列即可保持性能,适合大规模跨队列分析
宏基因组疾病预测通常依赖于来自庞大但不完整的参考目录的物种丰度表,限制了分辨率并丢弃了读段中的宝贵信息。为克服这些局限,我们提出MetagenBERT,一个基于Transformer的框架,可直接从原始DNA序列生成端到端的宏基因组嵌入,无需分类或功能注释。读段通过基础基因组语言模型(DNABERT2和专用于微生物组的DNABERTMS)进行嵌入,再通过基于FAISS加速的KMeans聚类策略聚合。每个宏基因组被表示为聚类丰度向量,总结其嵌入读段的分布。我们在五个基准肠道菌群数据集(肝硬化、2型糖尿病、肥胖、IBD、结直肠癌)上评估该方法,MetagenBERT在大多数任务中达到竞争性或更优的AUC表现。结合两种表示进一步提升预测效果,表明分类与嵌入信号具有互补性。即使仅使用10%的读段,聚类仍保持鲁棒性,凸显宏基因组中存在大量冗余,可实现显著计算节省。我们还引入MetagenBERT Glob Mcardis,一个在大型、表型多样的MetaCardis队列上训练的跨队列变体,可迁移到其他数据集并保留对未见表型的预测能力,表明构建宏基因组基础模型的可行性。稳健性分析(PERMANOVA、PERMDISP、熵)显示不同状态在子样本间始终一致分离。总体而言,MetagenBERT提供了一种可扩展、无注释的宏基因组表示,指向未来在异质队列和测序技术下实现表型感知的泛化。
原文摘要 · Abstract (English)
Metagenomic disease prediction commonly relies on species abundance tables derived from large, incomplete reference catalogs, constraining resolution and discarding valuable information contained in DNA reads. To overcome these limitations, we introduce MetagenBERT, a Transformer based framework that produces end to end metagenome embeddings directly from raw DNA sequences, without taxonomic or functional annotations. Reads are embedded using foundational genomic language models (DNABERT2 and the microbiome specialized DNABERTMS), then aggregated through a scalable clustering strategy based on FAISS accelerated KMeans. Each metagenome is represented as a cluster abundance vector summarizing the distribution of its embedded reads. We evaluate this approach on five benchmark gut microbiome datasets (Cirrhosis, T2D, Obesity, IBD, CRC). MetagenBERT achieves competitive or superior AUC performance relative to species abundance baselines across most tasks. Concatenating both representations further improves prediction, demonstrating complementarity between taxonomic and embedding derived signals. Clustering remains robust when applied to as little as 10% of reads, highlighting substantial redundancy in metagenomes and enabling major computational gains. We additionally introduce MetagenBERT Glob Mcardis, a cross cohort variant trained on the large, phenotypically diverse MetaCardis cohort and transferred to other datasets, retaining predictive signal including for unseen phenotypes, indicating the feasibility of a foundation model for metagenome representation. Robustness analyses (PERMANOVA, PERMDISP, entropy) show consistent separation of different states across subsamples. Overall, MetagenBERT provides a scalable, annotation free representation of metagenomes pointing toward future phenotype aware generalization across heterogeneous cohorts and sequencing technologies.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。