arXiv:2602.17680cs.LG2026-02被引 2

将蛋白语言模型与通用大模型结合,提升生物推理能力

BioBridge: Bridging Proteins and Language for Enhanced Biological Reasoning with LLMs

  • 通过增量持续预训练融合蛋白知识与通用语义
  • 在EC、BindingDB等数据集上表现媲美主流蛋白模型
  • 适合需要跨领域生物推理的研究者使用

现有蛋白语言模型(PLMs)适应多任务能力有限,跨生物场景泛化性差;通用大语言模型(LLMs)缺乏对蛋白序列的解析能力与领域知识,难以实现有效生物语义推理。为此,我们提出BioBridge,一种面向蛋白质理解的领域自适应持续预训练框架。该框架采用领域增量持续预训练(DICP),同步注入蛋白领域知识与通用推理语料,有效缓解灾难性遗忘。通过PLM-投影器-LLM跨模态对齐管道,将蛋白序列嵌入映射至语言模型语义空间。最终采用端到端优化,统一支持蛋白性质预测与知识问答等多种任务。BioBridge在多个蛋白基准测试(如EC、BindingDB)上的表现与主流PLMs相当,在通用理解任务(如MMLU、RACE)上也达到与LLMs相当的水平,展现出结合领域专长与通用语言能力的创新优势。

原文摘要 · Abstract (English)

Existing Protein Language Models (PLMs) often suffer from limited adaptability to multiple tasks and exhibit poor generalization across diverse biological contexts. In contrast, general-purpose Large Language Models (LLMs) lack the capability to interpret protein sequences and fall short in domain-specific knowledge, limiting their capacity for effective biosemantic reasoning. To combine the advantages of both, we propose BioBridge, a domain-adaptive continual pretraining framework for protein understanding. This framework employs Domain-Incremental Continual Pre-training (DICP) to infuse protein domain knowledge and general reasoning corpus into a LLM simultaneously, effectively mitigating catastrophic forgetting. Cross-modal alignment is achieved via a PLM-Projector-LLM pipeline, which maps protein sequence embeddings into the semantic space of the language model. Ultimately, an end-to-end optimization is adopted to uniformly support various tasks, including protein property prediction and knowledge question-answering. Our proposed BioBridge demonstrates performance comparable to that of mainstream PLMs on multiple protein benchmarks, such as EC and BindingDB. It also achieves results on par with LLMs on general understanding tasks like MMLU and RACE. This showcases its innovative advantage of combining domain-specific adaptability with general-purpose language competency.

蛋白质建模大模型跨模态

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。