arXiv:2410.11499q-bio.GNcs.AI2024-10被引 1

小模型实现多模态生物序列理解,性能媲美大模型

BSM: Small but Powerful Biological Sequence Model for Genes and Proteins

  • 用混合数据训练,学习基因与蛋白间的跨模态关系
  • 110M参数模型在单/多模态任务上表现接近更大模型
  • 首次展示多模态上下文学习能力,适合生物序列分析者

DNA、RNA和蛋白质等生物序列建模对理解基因调控和蛋白质合成至关重要。现有模型多局限于单一类型或分开处理,难以捕捉跨模态关联。本文提出BSM,一个仅110M参数的小型多模态生物序列基础模型,训练数据包括RefSeq、基因相关序列及网络上交错的生物序列,分别覆盖基因流、基因-蛋白关系和自然共现数据。通过混合模态训练,BSM显著提升学习效率与跨模态表征能力,在单模态和多模态任务中表现媲美更大模型,并首次展现出多模态上下文学习能力。进一步扩展至270M参数后性能持续提升,验证了其在多模态生物序列建模中的重要潜力。

原文摘要 · Abstract (English)

Modeling biological sequences such as DNA, RNA, and proteins is crucial for understanding complex processes like gene regulation and protein synthesis. However, most current models either focus on a single type or treat multiple types of data separately, limiting their ability to capture cross-modal relationships. We propose that by learning the relationships between these modalities, the model can enhance its understanding of each type. To address this, we introduce BSM, a small but powerful mixed-modal biological sequence foundation model, trained on three types of data: RefSeq, Gene Related Sequences, and interleaved biological sequences from the web. These datasets capture the genetic flow, gene-protein relationships, and the natural co-occurrence of diverse biological data, respectively. By training on mixed-modal data, BSM significantly enhances learning efficiency and cross-modal representation, outperforming models trained solely on unimodal data. With only 110M parameters, BSM achieves performance comparable to much larger models across both single-modal and mixed-modal tasks, and uniquely demonstrates in-context learning capability for mixed-modal tasks, which is absent in existing models. Further scaling to 270M parameters demonstrates even greater performance gains, highlighting the potential of BSM as a significant advancement in multimodal biological sequence modeling.

生物序列多模态小模型基因蛋白

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。