arXiv:2411.15418q-bio.BMcs.LG2024-11被引 6

用向量方法实现百亿分子对全蛋白组的快速药物靶点预测

Scaling Structure Aware Virtual Screening to Billions of Molecules with SPRINT

  • 基于自注意力结构和结构感知语言模型,学习药物-靶点联合嵌入
  • 在多个基准上达到顶尖富集效果,百万级搜索仅需16分钟
  • 支持残基级注意力解释,适合药物重定位与机制发现研究

针对蛋白靶标的化合物虚拟筛选可加速药物研发,但基于结构的方法(如分子对接)速度过慢,难以开展全蛋白组范围筛查,限制了其在脱靶效应或新作用机制发现中的应用。近期,基于向量的蛋白语言模型方法成为替代方案,避免了显式三维结构建模。本文提出SPRINT,一种面向全化学库与全蛋白组的向量化虚拟筛选方法,用于预测药物-靶点相互作用及新作用机制。SPRINT采用自注意力架构与结构感知蛋白语言模型,学习药物-靶点共嵌入以实现结合物预测、搜索与检索。在LIT-PCBA、DTI分类及结合亲和力预测等多个基准上,SPRINT均达到最先进性能,并提供残基级别的注意力热图以增强可解释性。此外,SPRINT具有极快速度:对整个人类蛋白组与67亿分子的ENAMINE Real数据库进行查询,每蛋白返回前100个高可能性结合物,仅需16分钟。SPRINT有望实现前所未有的虚拟筛选规模,为计算机辅助药物重定位与开发开辟新路径。项目已上线ColabScreen:https://bit.ly/colab-screen

原文摘要 · Abstract (English)

Virtual screening of small molecules against protein targets can accelerate drug discovery and development by predicting drug-target interactions (DTIs). However, structure-based methods like molecular docking are too slow to allow for broad proteome-scale screens, limiting their application in screening for off-target effects or new molecular mechanisms. Recently, vector-based methods using protein language models (PLMs) have emerged as a complementary approach that bypasses explicit 3D structure modeling. Here, we develop SPRINT, a vector-based approach for screening entire chemical libraries against whole proteomes for DTIs and novel mechanisms of action. SPRINT improves on prior work by using a self-attention based architecture and structure-aware PLMs to learn drug-target co-embeddings for binder prediction, search, and retrieval. SPRINT achieves SOTA enrichment factors in virtual screening on LIT-PCBA, DTI classification benchmarks, and binding affinity prediction benchmarks, while providing interpretability in the form of residue-level attention maps. In addition to being both accurate and interpretable, SPRINT is ultra-fast: querying the whole human proteome against the ENAMINE Real Database (6.7B drugs) for the 100 most likely binders per protein takes 16 minutes. SPRINT promises to enable virtual screening at an unprecedented scale, opening up new opportunities for in silico drug repurposing and development. SPRINT is available on the web as ColabScreen: https://bit.ly/colab-screen

虚拟筛选药物重定位蛋白语言模型快速预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。