MAMMAL统一建模蛋白、小分子与组学数据,提升药物发现效率。
MAMMAL -- Molecular Aligned Multi-Modal Architecture and Language
- 构建跨模态对齐架构,统一处理蛋白质、小分子与组学数据
- 在11项任务中9项达新SOTA,3个抗原复合物分类性能更优
- 适合生物医学研究者用于多模态药物设计与机制探索
将大型语言模型应用于海量生物数据,有望揭示疾病机制并加速药物研发。然而,现有模型常局限于单一模态(如小分子、蛋白质或转录组数据),难以捕捉复杂的多模态交互。有效药物发现需要能整合多种生物实体并支持预测与生成的计算工具,而当前模型难以应对这一挑战。为此,我们提出MAMMAL——分子对齐多模态架构与语言模型,一种适用于大规模跨模态生物数据(包括蛋白质、小分子和组学数据)的多任务基础模型。MAMMAL采用结构化提示语法,支持分类、回归与生成任务,可处理标记与标量输入输出。在11项下游任务上评估,其在9项中达到新SOTA,2项表现接近SOTA,且所有任务均在统一架构下完成,优于以往任务专用模型。此外,在阿尔法折叠3抗体-抗原及纳米抗体-抗原复合物结合预测中,MAMMAL在4个目标中的3个表现出显著更优的分类性能。模型代码与预训练权重已公开于https://github.com/BiomedSciAI/biomed-multi-alignment 和 https://huggingface.co/ibm/biomed.omics.bl.sm.ma-ted-458m。
原文摘要 · Abstract (English)
Large language models applied to vast biological datasets have the potential to transform biology by uncovering disease mechanisms and accelerating drug development. However, current models are often siloed, trained separately on small-molecules, proteins, or transcriptomic data, limiting their ability to capture complex, multi-modal interactions. Effective drug discovery requires computational tools that integrate multiple biological entities while supporting prediction and generation, a challenge existing models struggle to address. For this purpose, we present MAMMAL - Molecular Aligned Multi-Modal Architecture and Language - a versatile method applied to create a multi-task foundation model that learns from large-scale biological datasets across diverse modalities, including proteins, small-molecules, and omics. MAMMAL's structured prompt syntax supports classification, regression, and generation tasks while handling token and scalar inputs and outputs. Evaluated on eleven diverse downstream tasks, it reaches a new state of the art (SOTA) in nine tasks and is comparable to SOTA in two tasks, all within a unified architecture, unlike prior task-specific models. Additionally, we explored Alphafold 3 binding prediction capabilities on antibody-antigen and nanobody-antigen complexes showing significantly better classification performance of MAMMAL in 3 out of 4 targets. The model code and pretrained weights are publicly available at https://github.com/BiomedSciAI/biomed-multi-alignment and https://huggingface.co/ibm/biomed.omics.bl.sm.ma-ted-458m
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。