arXiv:2409.13057q-bio.QMcs.CL2024-09被引 4

用自然语言处理预测蛋白-配体相互作用,加速药物研发

Natural Language Processing Methods for the Study of Protein-Ligand Interactions

  • 将NLP技术应用于蛋白质和配体序列的建模与交互预测
  • 利用注意力机制和Transformer架构提升预测准确性
  • 适合药物发现与蛋白质工程领域的研究人员参考

自然语言处理(NLP)的最新进展激发了对蛋白-配体相互作用(PLIs)预测方法的研究兴趣,这在药物发现和蛋白质工程中具有重要意义,且面对日益增长的生物化学序列与结构数据。蛋白质和配体的表示方式与人类语言存在类比关系,使得NLP机器学习方法可用于推进PLI研究。本文综述了近期文献中NLP方法的应用场景与实现方式,讨论了长短期记忆网络、Transformer及注意力机制等关键机制。最后,我们总结了当前NLP方法在PLI研究中的局限性,并指出未来需解决的关键挑战。

原文摘要 · Abstract (English)

Recent advances in Natural Language Processing (NLP) have ignited interest in developing effective methods for predicting protein-ligand interactions (PLIs) given their relevance to drug discovery and protein engineering efforts and the ever-growing volume of biochemical sequence and structural data available. The parallels between human languages and the "languages" used to represent proteins and ligands have enabled the use of NLP machine learning approaches to advance PLI studies. In this review, we explain where and how such approaches have been applied in the recent literature and discuss useful mechanisms such as long short-term memory, transformers, and attention. We conclude with a discussion of the current limitations of NLP methods for the study of PLIs as well as key challenges that need to be addressed in future work.

蛋白-配体自然语言处理药物发现

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。