arXiv:2501.02045q-bio.GNcs.AI2025-01被引 16

用1.5万亿碱基废水数据训练70亿参数模型,监测疫情更早更准。

METAGENE-1: Metagenomic Foundation Model for Pandemic Monitoring

  • 基于废水基因组数据构建自回归Transformer模型
  • 在病原体检测与序列嵌入任务上达当前最佳表现
  • 适合公共卫生、流行病预警等场景使用

我们预训练了METAGENE-1,一个70亿参数的自回归Transformer模型,作为宏基因组基础模型,在包含超过1.5万亿碱基对的多样化宏基因组DNA和RNA序列新语料库上进行训练。该数据集来自大量人类污水样本,通过深度宏基因组(下一代)测序方法处理并测序。不同于仅关注单个基因组或特定物种的基因组模型,METAGENE-1旨在捕捉此类污水中全部基因组信息分布,以支持疫情监测与病原体检测任务。我们针对宏基因组序列设计了字节对编码(BPE)分词策略,并在此基础上完成模型预训练。本文首先详述预训练数据集、分词方案及模型架构,强调实现宏基因组数据有效建模的设计考量。随后展示模型在宏基因组数据上的预训练结果,包括损失值、系统指标及训练稳定性。最后,我们验证METAGENE-1在一组基因组基准测试和聚焦于人源病原体检测与基因组序列嵌入的新评估任务中的表现,结果显示其达到当前最优水平,展现出在疫情监测、生物监视和新兴健康威胁早期发现中的应用潜力。

原文摘要 · Abstract (English)

We pretrain METAGENE-1, a 7-billion-parameter autoregressive transformer model, which we refer to as a metagenomic foundation model, on a novel corpus of diverse metagenomic DNA and RNA sequences comprising over 1.5 trillion base pairs. This dataset is sourced from a large collection of human wastewater samples, processed and sequenced using deep metagenomic (next-generation) sequencing methods. Unlike genomic models that focus on individual genomes or curated sets of specific species, the aim of METAGENE-1 is to capture the full distribution of genomic information present within this wastewater, to aid in tasks relevant to pandemic monitoring and pathogen detection. We carry out byte-pair encoding (BPE) tokenization on our dataset, tailored for metagenomic sequences, and then pretrain our model. In this paper, we first detail the pretraining dataset, tokenization strategy, and model architecture, highlighting the considerations and design choices that enable the effective modeling of metagenomic data. We then show results of pretraining this model on our metagenomic dataset, providing details about our losses, system metrics, and training stability over the course of pretraining. Finally, we demonstrate the performance of METAGENE-1, which achieves state-of-the-art results on a set of genomic benchmarks and new evaluations focused on human-pathogen detection and genomic sequence embedding, showcasing its potential for public health applications in pandemic monitoring, biosurveillance, and early detection of emerging health threats.

宏基因组疫情监测基础模型废水测序

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。