GENErator是能处理9.8万碱基对的基因组生成模型,可零样本预测突变影响并设计功能序列。
GENERator: A Long-Context Generative Genomic Foundation Model
- 基于3860亿碱基训练,支持98k长序列生成与建模
- 无需微调即实现媲美或超越现有模型的生成准确率和效率
- 可零样本预测突变效应,还能设计具有目标活性的增强子序列
DNA测序技术快速发展带来了海量基因组数据,但解读与工程化基因组功能仍是根本挑战。现有大语言模型在基因组分析中受限于训练范围窄、生成能力弱或计算成本高。我们提出GENErator,一种面向长序列(98,000碱基)的生成式基因组基础模型,基于3860亿个真核生物碱基预训练。无需任务微调,GENErator展现强内在能力:无监督嵌入分析揭示系统发育一致结构;序列恢复基准测试显示生成精度达到或超过当前最优模型,且计算效率显著提升。在零样本设定下,其突变效应预测性能可比肩基于比对的方法,同时完全免于比对,跨物种通用。经任务微调后,在主流基因组基准上达到领先水平。我们进一步展示实际生成应用:能生成可翻译为结构合理蛋白的编码序列,并通过提示引导框架设计具有目标活性谱的顺式调控元件,包括经高通量UMI-STARR-seq验证的合成超增强子。这些结果确立了GENErator作为高效且生物学可信的基因组解析与可编程序列设计框架。代码与补充资源见https://github.com/GenerTeam/GENERator。
原文摘要 · Abstract (English)
The rapid advancement of DNA sequencing has produced vast genomic datasets, yet interpreting and engineering genomic function remain fundamental challenges. Recent large language models have opened new avenues for genomic analysis, but existing approaches are often limited by restricted training scope, constrained generative capability, or prohibitive computational cost. We introduce GENErator, a generative genomic foundation model for long-context DNA modeling, with a context length of 98k nucleotides, pre-trained on 386 billion nucleotides of eukaryotic DNA. Without task-specific fine-tuning, GENERator exhibits strong intrinsic capabilities: unsupervised embedding analyses reveal phylogenetically coherent structure, and sequence recovery benchmarks demonstrate generative accuracy comparable to or exceeding state-of-the-art models with substantially improved computational efficiency. In a zero-shot setting, GENERator achieves competitive variant effect prediction performance relative to alignment-based methods, while remaining fully alignment-free and broadly applicable across species. With task-specific fine-tuning, the model attains leading performance on established genomic benchmarks. We further demonstrate practical generative applications. GENERator can generate protein-coding DNA sequences that translate into structurally plausible proteins and, through a prompt-guided design framework, design cis-regulatory elements with targeted activity profiles, including synthetic super-enhancers validated by high-throughput UMI-STARR-seq assays. Together, these results establish GENERator as an efficient and biologically grounded framework for genomic interpretation and programmable sequence design. Code and supplementary resources are available at https://github.com/GenerTeam/GENERator.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。