arXiv:2411.00004q-bio.BMcs.LG2024-11

用Transformer加速蛋白质分子对接,百倍提速且精度不降。

RapidDock: Unlocking Proteome-scale Molecular Docking

  • 基于Transformer设计,引入3D相对距离嵌入和对称不变损失函数。
  • 在Posebusters和DockGen上分别达到52.1%和44.0%的准确率(RMSD<2Å)。
  • 单卡推理仅需0.04秒,适合全蛋白组规模药物筛选。

加速分子对接——预测分子与蛋白靶点结合方式——可推动小分子药物研发并变革医学。然而,现有对接工具速度太慢,难以对所有相关蛋白进行药物候选物筛选,常导致遗漏药候选或临床试验中出现意外副作用。为填补这一空白,我们提出RapidDock,一种高效的基于Transformer的盲对接模型。RapidDock相比现有方法至少提速100倍,且不牺牲精度。在Posebusters和DockGen基准测试中,成功率达52.1%和44.0%(RMSD<2Å)。单次推理平均耗时仅0.04秒,凸显其在大规模对接研究中的潜力。我们分析了RapidDock的关键设计:将3D结构的相对距离嵌入注意力矩阵、在蛋白质折叠数据上预训练,以及采用对称性不变的自定义损失函数。

原文摘要 · Abstract (English)

Accelerating molecular docking -- the process of predicting how molecules bind to protein targets -- could boost small-molecule drug discovery and revolutionize medicine. Unfortunately, current molecular docking tools are too slow to screen potential drugs against all relevant proteins, which often results in missed drug candidates or unexpected side effects occurring in clinical trials. To address this gap, we introduce RapidDock, an efficient transformer-based model for blind molecular docking. RapidDock achieves at least a $100 \times$ speed advantage over existing methods without compromising accuracy. On the Posebusters and DockGen benchmarks, our method achieves $52.1\%$ and $44.0\%$ success rates ($\text{RMSD}<2$Å), respectively. The average inference time is $0.04$ seconds on a single GPU, highlighting RapidDock's potential for large-scale docking studies. We examine the key features of RapidDock that enable leveraging the transformer architecture for molecular docking, including the use of relative distance embeddings of $3$D structures in attention matrices, pre-training on protein folding, and a custom loss function invariant to molecular symmetries.

分子对接Transformer药物发现加速计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。