arXiv:2502.03499q-bio.GNcs.AI2025-02被引 7

一个能跨模态多任务的基因组大模型,一次训练搞定多种基因分析。

Omni-DNA: A Unified Genomic Foundation Model for Cross-Modal and Multi-Task Learning

  • 用自回归预训练+多任务微调,统一处理不同基因任务。
  • 在26个任务中18个达到顶尖水平,10个甲基化乙酰化任务同时超越单任务模型。
  • 可将DNA序列转为文字描述或图像,适合复杂基因功能研究者。

大型语言模型在多种任务中表现出色,但现有基因组基础模型(GFMs)仍需为每个下游任务单独微调,随着模型规模增大带来显著开销。且现有GFMs受限于固定输出格式,难以适配多样基因任务。本文重新审视基于Transformer的自回归模型,提出Omni-DNA,一组参数规模从2000万到10亿不等的跨模态多任务模型。方法分为两阶段:(i) 在DNA序列上以预测下一个词为目标进行预训练;(ii) 扩展多模态任务特定标记并同时微调多个下游任务。在Nucleotide Transformer和GB基准测试中,Omni-DNA在26个任务中的18个达到当前最优性能。通过多任务微调,一次性完成10个乙酰化与甲基化任务,优于分别训练的模型。最后,设计两个复杂任务:DNA2Function(将DNA序列映射为文本功能描述)和Needle-in-DNA(映射为图像),验证了Omni-DNA的跨模态能力,拓展了基因组应用边界。所有模型已公开于https://huggingface.co/collections/zehui127。

原文摘要 · Abstract (English)

Large Language Models (LLMs) demonstrate remarkable generalizability across diverse tasks, yet genomic foundation models (GFMs) still require separate finetuning for each downstream application, creating significant overhead as model sizes grow. Moreover, existing GFMs are constrained by rigid output formats, limiting their applicability to various genomic tasks. In this work, we revisit the transformer-based auto-regressive models and introduce Omni-DNA, a family of cross-modal multi-task models ranging from 20 million to 1 billion parameters. Our approach consists of two stages: (i) pretraining on DNA sequences with next token prediction objective, and (ii) expanding the multi-modal task-specific tokens and finetuning for multiple downstream tasks simultaneously. When evaluated on the Nucleotide Transformer and GB benchmarks, Omni-DNA achieves state-of-the-art performance on 18 out of 26 tasks. Through multi-task finetuning, Omni-DNA addresses 10 acetylation and methylation tasks at once, surpassing models trained on each task individually. Finally, we design two complex genomic tasks, DNA2Function and Needle-in-DNA, which map DNA sequences to textual functional descriptions and images, respectively, indicating Omni-DNA's cross-modal capabilities to broaden the scope of genomic applications. All the models are available through https://huggingface.co/collections/zehui127

基因组模型多任务学习跨模态自回归

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。