arXiv:2409.17726cs.LG2024-09被引 2

用可解释的蛋白结构表示法,提升蛋白质预测与功能分析的透明度。

Recent advances in interpretable machine learning using structure-based protein representations

  • 基于多分辨率蛋白3D结构表示,构建可解释的机器学习模型
  • 实现蛋白质结构、功能及相互作用的高可信度预测
  • 适合生物信息学与药物设计研究者快速理解模型决策

机器学习的最新进展正在重塑结构生物学领域。例如,突破性的蛋白质结构预测神经网络AlphaFold,已广泛被研究人员采用。其易于使用的界面和可解释的输出结果(如用于着色预测结构的置信度分数),使非机器学习专家也能轻松使用。本文综述了从低到高分辨率的多种蛋白3D结构表示方法,并展示了可解释机器学习如何支持蛋白质结构预测、蛋白质功能推断及蛋白质-蛋白质相互作用分析。本综述还强调了对基于结构的蛋白表示进行解释与可视化的重要性,有助于提升模型可解释性并促进知识发现。发展此类可解释方法有望进一步加速药物研发与蛋白质设计等领域的发展。

原文摘要 · Abstract (English)

Recent advancements in machine learning (ML) are transforming the field of structural biology. For example, AlphaFold, a groundbreaking neural network for protein structure prediction, has been widely adopted by researchers. The availability of easy-to-use interfaces and interpretable outcomes from the neural network architecture, such as the confidence scores used to color the predicted structures, have made AlphaFold accessible even to non-ML experts. In this paper, we present various methods for representing protein 3D structures from low- to high-resolution, and show how interpretable ML methods can support tasks such as predicting protein structures, protein function, and protein-protein interactions. This survey also emphasizes the significance of interpreting and visualizing ML-based inference for structure-based protein representations that enhance interpretability and knowledge discovery. Developing such interpretable approaches promises to further accelerate fields including drug development and protein design.

可解释AI蛋白质结构生物信息学机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。