将ESM2模型输入长度扩展至2048个氨基酸,支持长蛋白序列分析。
Scaling Up ESM2 Architectures for Long Protein Sequences Analysis: Long and Quantized Approaches
- 通过改进架构实现输入长度翻倍,突破原1022氨基酸限制。
- 新版本在长序列数据上表现更优,减少分段处理带来的信息损失。
- 适合研究长链蛋白结构与功能的生物学家和计算生物学工作者。
Transformer架构在自然语言处理中取得卓越成果,推动其在生物序列分析中的应用,尤其在蛋白质序列领域。其中,基于数十亿蛋白质预训练的ESM2架构已成为该领域的主流方法。然而,原始ESM2模型存在输入长度限制,最大仅支持1,022个氨基酸,导致长于该长度的序列需经分割等预处理操作,可能丢失关键信息。本文提出长序列版与量化版ESM2架构,将输入上限提升至2,048个氨基酸,显著增强对长蛋白序列的建模能力,为全长度蛋白分析提供更可靠的工具。
原文摘要 · Abstract (English)
Various approaches utilizing Transformer architectures have achieved state-of-the-art results in Natural Language Processing (NLP). Based on this success, numerous architectures have been proposed for other types of data, such as in biology, particularly for protein sequences. Notably among these are the ESM2 architectures, pre-trained on billions of proteins, which form the basis of various state-of-the-art approaches in the field. However, the ESM2 architectures have a limitation regarding input size, restricting it to 1,022 amino acids, which necessitates the use of preprocessing techniques to handle sequences longer than this limit. In this paper, we present the long and quantized versions of the ESM2 architectures, doubling the input size limit to 2,048 amino acids.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。