大语言模型让蛋白质科学实现跨任务通用推理,推动结构、功能与设计研究突破。
Computational Protein Science in the Era of Large Language Models (LLMs)
- 基于大语言模型构建蛋白质语言模型,融合序列、结构与功能知识
- 在蛋白质结构预测、功能识别与设计中取得显著进展,覆盖抗体、酶与药物研发
- 为跨领域蛋白质研究提供通用工具,适合生物信息与药物研发人员参考
蛋白质在生命科学中至关重要,计算蛋白质科学致力于揭示蛋白质序列-结构-功能关系并推动应用。过去几十年,人工智能在特定建模任务中取得显著成果,但存在难以理解蛋白质序列语义、泛化能力不足等问题。近期大语言模型(LLMs)凭借卓越的语言处理与泛化能力,成为AI里程碑,推动多领域综合进步。研究人员将LLM技术引入蛋白质科学,发展出蛋白质语言模型(pLMs),能有效捕捉蛋白质基础知识,并广泛应用于序列-结构-功能推理任务。本文系统综述了pLMs的发展:按掌握的知识类型分为底层序列模式、显式结构功能信息及外部科学语言三类;总结其在蛋白质结构预测、功能预测与设计中的应用成果;展示其在抗体设计、酶设计与药物发现中的实践价值;最后展望该快速发展的未来方向。
原文摘要 · Abstract (English)
Considering the significance of proteins, computational protein science has always been a critical scientific field, dedicated to revealing knowledge and developing applications within the protein sequence-structure-function paradigm. In the last few decades, Artificial Intelligence (AI) has made significant impacts in computational protein science, leading to notable successes in specific protein modeling tasks. However, those previous AI models still meet limitations, such as the difficulty in comprehending the semantics of protein sequences, and the inability to generalize across a wide range of protein modeling tasks. Recently, LLMs have emerged as a milestone in AI due to their unprecedented language processing & generalization capability. They can promote comprehensive progress in fields rather than solving individual tasks. As a result, researchers have actively introduced LLM techniques in computational protein science, developing protein Language Models (pLMs) that skillfully grasp the foundational knowledge of proteins and can be effectively generalized to solve a diversity of sequence-structure-function reasoning problems. While witnessing prosperous developments, it's necessary to present a systematic overview of computational protein science empowered by LLM techniques. First, we summarize existing pLMs into categories based on their mastered protein knowledge, i.e., underlying sequence patterns, explicit structural and functional information, and external scientific languages. Second, we introduce the utilization and adaptation of pLMs, highlighting their remarkable achievements in promoting protein structure prediction, protein function prediction, and protein design studies. Then, we describe the practical application of pLMs in antibody design, enzyme design, and drug discovery. Finally, we specifically discuss the promising future directions in this fast-growing field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。