Gene42可处理19万碱基长基因组序列,突破现有模型极限。
Gene42: Long-Range Genomic Foundation Model With Dense Attention
- 采用密集自注意力机制的解码器架构,逐步扩展上下文至19.2万碱基。
- 在多种基因组任务上表现领先,如调控区域识别和变异致病性预测。
- 适合需要长距离基因组依赖分析的研究者使用。
我们提出Gene42,一类新型基因组基础模型(GFMs),可处理长达192,000个碱基对(bp)的基因组序列,并保持单碱基分辨率。该模型采用类似LLaMA的解码器架构,配备密集自注意力机制。初始在4,096 bp固定长度序列上训练,随后通过持续预训练将上下文长度扩展至192,000 bp。这一迭代扩展使模型能全面处理大规模基因组数据,捕捉人类基因组中复杂的模式与依赖关系。Gene42是首个能在基因组领域处理如此长上下文的密集注意力模型,挑战了以往依赖卷积等机制的状态空间模型。预训练模型展现出极低困惑度与高重建准确率,表明其强大的基因组建模能力。在多个基因组基准测试中,其在生物类型分类、调控区域识别、染色质特征预测、变异致病性预测及物种分类等任务上均达到当前最优性能。模型已公开于huggingface.co/inceptionai。
原文摘要 · Abstract (English)
We introduce Gene42, a novel family of Genomic Foundation Models (GFMs) designed to manage context lengths of up to 192,000 base pairs (bp) at a single-nucleotide resolution. Gene42 models utilize a decoder-only (LLaMA-style) architecture with a dense self-attention mechanism. Initially trained on fixed-length sequences of 4,096 bp, our models underwent continuous pretraining to extend the context length to 192,000 bp. This iterative extension allowed for the comprehensive processing of large-scale genomic data and the capture of intricate patterns and dependencies within the human genome. Gene42 is the first dense attention model capable of handling such extensive long context lengths in genomics, challenging state-space models that often rely on convolutional operators among other mechanisms. Our pretrained models exhibit notably low perplexity values and high reconstruction accuracy, highlighting their strong ability to model genomic data. Extensive experiments on various genomic benchmarks have demonstrated state-of-the-art performance across multiple tasks, including biotype classification, regulatory region identification, chromatin profiling prediction, variant pathogenicity prediction, and species classification. The models are publicly available at huggingface.co/inceptionai.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。