系统梳理蛋白质表征学习方法与应用,助力药物研发与生命研究。
Advances in Protein Representation Learning: Methods, Applications, and Future Directions
- 按特征、序列、结构等五类方法分类梳理蛋白表征技术。
- 整合常用数据库,支持模型训练与评估的资源建设。
- 覆盖药物发现等多领域应用,展望未来技术挑战与方向。
蛋白质是复杂的生物分子,在分子生物学、医学研究和药物发现中具有核心地位。解析其多层次结构与多样功能,对深化生命分子层面理解至关重要。蛋白质表征学习(PRL)作为一项变革性方法,能够从蛋白数据中提取有意义的计算表示,以应对这些挑战。本文全面综述了PRL研究,将方法分为五大类:基于特征、基于序列、基于结构、多模态及复杂系统方法。为支持该快速发展的领域,我们介绍了广泛使用的蛋白序列、结构与功能数据库,为模型开发与评估提供关键资源。此外,探讨了这些方法在多个领域的多样化应用,展示其广泛影响。最后,讨论当前关键技术挑战,并提出未来发展方向,为持续创新提供启示。
原文摘要 · Abstract (English)
Proteins are complex biomolecules that play a central role in various biological processes, making them critical targets for breakthroughs in molecular biology, medical research, and drug discovery. Deciphering their intricate, hierarchical structures, and diverse functions is essential for advancing our understanding of life at the molecular level. Protein Representation Learning (PRL) has emerged as a transformative approach, enabling the extraction of meaningful computational representations from protein data to address these challenges. In this paper, we provide a comprehensive review of PRL research, categorizing methodologies into five key areas: feature-based, sequence-based, structure-based, multimodal, and complex-based approaches. To support researchers in this rapidly evolving field, we introduce widely used databases for protein sequences, structures, and functions, which serve as essential resources for model development and evaluation. We also explore the diverse applications of these approaches in multiple domains, demonstrating their broad impact. Finally, we discuss pressing technical challenges and outline future directions to advance PRL, offering insights to inspire continued innovation in this foundational field.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。