arXiv:2410.03553cs.CLq-bio.BM2024-10KDD被引 8

用结构增强的指令微调框架,让大模型读懂蛋白质功能。

Structure-Enhanced Protein Instruction Tuning: Towards General-Purpose Protein Understanding with LLMs

  • 在蛋白语言模型中加入结构感知模块,融合三维结构知识。
  • 构建最大蛋白指令数据集,实现通用蛋白理解能力。
  • 采用专家混合机制,在不增加参数量下提升复杂功能预测。

蛋白质是生物过程中的核心分子,准确预测其性质与功能对生物学应用至关重要。尽管基于监督微调的蛋白语言模型(pLMs)已取得进展,但现有模型多针对特定任务定制,难以实现通用蛋白理解。本文提出结构增强型蛋白指令微调框架(SEPIT),通过引入结构感知模块增强pLMs的结构知识,并将其与大语言模型(LLMs)结合,推动通用蛋白理解。我们设计了一套新的指令微调流程:首先利用对比学习与结构去噪预热增强pLMs;其次通过图文描述指令建立基础蛋白理解;最后采用专家混合(MoEs)机制,在保持激活参数数量不变的前提下,捕捉更复杂的性质与功能信息。此外,我们构建了迄今最大最全面的蛋白指令数据集,支持通用蛋白理解模型的训练与评估。大量实验表明,SEPIT在开放生成和闭集问答任务上均优于闭源通用LLMs及基于蛋白知识训练的开源LLMs。

原文摘要 · Abstract (English)

Proteins, as essential biomolecules, play a central role in biological processes, including metabolic reactions and DNA replication. Accurate prediction of their properties and functions is crucial in biological applications. Recent development of protein language models (pLMs) with supervised fine tuning provides a promising solution to this problem. However, the fine-tuned model is tailored for particular downstream prediction task, and achieving general-purpose protein understanding remains a challenge. In this paper, we introduce Structure-Enhanced Protein Instruction Tuning (SEPIT) framework to bridge this gap. Our approach incorporates a novel structure-aware module into pLMs to enrich their structural knowledge, and subsequently integrates these enhanced pLMs with large language models (LLMs) to advance protein understanding. In this framework, we propose a novel instruction tuning pipeline. First, we warm up the enhanced pLMs using contrastive learning and structure denoising. Then, caption-based instructions are used to establish a basic understanding of proteins. Finally, we refine this understanding by employing a mixture of experts (MoEs) to capture more complex properties and functional information with the same number of activated parameters. Moreover, we construct the largest and most comprehensive protein instruction dataset to date, which allows us to train and evaluate the general-purpose protein understanding model. Extensive experiments on both open-ended generation and closed-set answer tasks demonstrate the superior performance of SEPIT over both closed-source general LLMs and open-source LLMs trained with protein knowledge.

蛋白语言模型指令微调结构感知大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。