首份系统综述蛋白大模型,覆盖架构、数据与应用全链条。
Protein Large Language Models: A Comprehensive Survey
- 构建蛋白大模型分类体系,梳理100+论文方法脉络。
- 揭示大规模序列数据如何提升结构预测与功能注释精度。
- 适合生物信息学与蛋白质工程研究者参考。
蛋白领域的大语言模型正推动蛋白质科学变革,实现更高效的结构预测、功能注释与设计。现有综述多聚焦特定方向,本文首次全面梳理蛋白大模型,涵盖其架构、训练数据集、评估指标及多样化应用。通过系统分析超过100篇文献,提出前沿蛋白大模型的结构化分类体系,分析其如何利用大规模蛋白序列数据提升准确性,并探讨其在蛋白质工程与生物医学研究中的潜力。同时讨论关键挑战与未来方向,强调蛋白大模型在蛋白质科学研究中作为核心工具的地位。相关资源已维护于 https://github.com/Yijia-Xiao/Protein-LLM-Survey。
原文摘要 · Abstract (English)
Protein-specific large language models (Protein LLMs) are revolutionizing protein science by enabling more efficient protein structure prediction, function annotation, and design. While existing surveys focus on specific aspects or applications, this work provides the first comprehensive overview of Protein LLMs, covering their architectures, training datasets, evaluation metrics, and diverse applications. Through a systematic analysis of over 100 articles, we propose a structured taxonomy of state-of-the-art Protein LLMs, analyze how they leverage large-scale protein sequence data for improved accuracy, and explore their potential in advancing protein engineering and biomedical research. Additionally, we discuss key challenges and future directions, positioning Protein LLMs as essential tools for scientific discovery in protein science. Resources are maintained at https://github.com/Yijia-Xiao/Protein-LLM-Survey.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。