arXiv:2507.13646cs.LGcs.AI2025-07综述被引 4

综述基于Transformer的蛋白序列分析与设计模型进展

A Comprehensive Review of Transformer-based language models for Protein Sequence Analysis and Design

  • 系统梳理蛋白序列分析中Transformer模型的应用思路
  • 涵盖功能预测、结构识别、新蛋白生成等多类任务
  • 适合生物信息学与AI交叉研究者参考

Transformer模型在自然语言处理领域影响深远,其成功也推动了在生物信息学等领域的应用。本文综述了近年来基于Transformer的模型在蛋白序列分析与设计中的进展。重点讨论了基因本体、蛋白功能与结构识别、从头蛋白生成及蛋白结合等应用方向。系统分析了相关工作的优势与局限,旨在为读者提供全面洞察。最后指出当前研究的不足,并探讨未来发展方向。本文有助于研究人员把握该领域的最新进展,指导后续研究。

原文摘要 · Abstract (English)

The impact of Transformer-based language models has been unprecedented in Natural Language Processing (NLP). The success of such models has also led to their adoption in other fields including bioinformatics. Taking this into account, this paper discusses recent advances in Transformer-based models for protein sequence analysis and design. In this review, we have discussed and analysed a significant number of works pertaining to such applications. These applications encompass gene ontology, functional and structural protein identification, generation of de novo proteins and binding of proteins. We attempt to shed light on the strength and weaknesses of the discussed works to provide a comprehensive insight to readers. Finally, we highlight shortcomings in existing research and explore potential avenues for future developments. We believe that this review will help researchers working in this field to have an overall idea of the state of the art in this field, and to orient their future studies.

蛋白设计Transformer生物信息学序列分析

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。