arXiv:2411.04150q-bio.QMcs.LG2024-11被引 1

用语言模型预测药物靶点结合亲和力,不依赖三维结构。

BAPULM: Binding Affinity Prediction using Language Models

  • 用ProtT5和MolFormer提取蛋白与配体的序列表征。
  • 在三个数据集上准确率最高达0.925,优于传统方法。
  • 适合做药物筛选的快速初筛,尤其无三维结构时。

识别药物-靶点相互作用对开发有效药物至关重要。结合亲和力量化此类相互作用,传统方法依赖计算密集的三维结构数据。相比之下,语言模型可高效处理序列数据,提供分子表征的新路径。本文提出BAPULM,一种基于序列的创新框架,利用ProtT5-XL-U50获取蛋白质的化学隐表示,通过MolFormer获取配体表征,无需复杂三维构型。该方法在基准数据集上验证,得分能力(R)分别为0.925±0.043(benchmark1k2101)、0.914±0.004(Test2016_290)和0.8132±0.001(CSAR-HiQ_36)。结果表明BAPULM在不同数据集上具有鲁棒性和高准确性,凸显序列模型在计算机辅助药物设计中的潜力,为无三维结构依赖的配体筛选提供可扩展替代方案。

原文摘要 · Abstract (English)

Identifying drug-target interactions is essential for developing effective therapeutics. Binding affinity quantifies these interactions, and traditional approaches rely on computationally intensive 3D structural data. In contrast, language models can efficiently process sequential data, offering an alternative approach to molecular representation. In the current study, we introduce BAPULM, an innovative sequence-based framework that leverages the chemical latent representations of proteins via ProtT5-XL-U50 and ligands through MolFormer, eliminating reliance on complex 3D configurations. Our approach was validated extensively on benchmark datasets, achieving scoring power (R) values of 0.925 $\pm$ 0.043, 0.914 $\pm$ 0.004, and 0.8132 $\pm$ 0.001 on benchmark1k2101, Test2016_290, and CSAR-HiQ_36, respectively. These findings indicate the robustness and accuracy of BAPULM across diverse datasets and underscore the potential of sequence-based models in-silico drug discovery, offering a scalable alternative to 3D-centric methods for screening potential ligands.

药物发现语言模型亲和力预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。