arXiv:2601.05792cs.LGcs.AI2026-01中稿 · the Generative and…被引 3

用多模态对比学习提升药物靶点相互作用预测精度

Tensor-DTI: Enhancing Biomolecular Interaction Prediction with Contrastive Embedding Learning

  • 融合分子图、蛋白语言模型和结合位点预测的多模态嵌入
  • 在多个基准上超越现有序列与图模型,百亿化合物库筛选有效
  • 适用于蛋白-RNA、肽-蛋白等交互预测,结果更可解释

准确预测药物-靶点相互作用(DTI)对计算药物发现至关重要,但现有模型常依赖单一模态的预定义分子描述符或序列嵌入,代表性有限。本文提出Tensor-DTI,一种基于对比学习的框架,整合分子图、蛋白语言模型及结合位点预测的多模态嵌入,以增强相互作用建模。该方法采用孪生双编码器结构,能同时捕捉化学与结构特征,并区分相互作用与非相互作用对。在多个DTI基准上的评估表明,Tensor-DTI优于现有的序列与图基模型。我们在百亿级化学库上对CDK2进行了大规模推理实验,即使训练时未包含CDK2,仍生成了化学上合理的命中分布。在与Glide对接和Boltz-2共折叠的富集研究中,该模型在CDK2上保持竞争力,并在严格家族保留分割下,降低了恢复中高亲和力配体所需筛选预算。我们还探索其在蛋白-RNA和肽-蛋白相互作用中的适用性。结果表明,结合多模态信息与对比目标,可显著提升预测准确性,并构建更具可解释性和可靠性感知的虚拟筛选模型。

原文摘要 · Abstract (English)

Accurate drug-target interaction (DTI) prediction is essential for computational drug discovery, yet existing models often rely on single-modality predefined molecular descriptors or sequence-based embeddings with limited representativeness. We propose Tensor-DTI, a contrastive learning framework that integrates multimodal embeddings from molecular graphs, protein language models, and binding-site predictions to improve interaction modeling. Tensor-DTI employs a siamese dual-encoder architecture, enabling it to capture both chemical and structural interaction features while distinguishing interacting from non-interacting pairs. Evaluations on multiple DTI benchmarks demonstrate that Tensor-DTI outperforms existing sequence-based and graph-based models. We also conduct large-scale inference experiments on CDK2 across billion-scale chemical libraries, where Tensor-DTI produces chemically plausible hit distributions even when CDK2 is withheld from training. In enrichment studies against Glide docking and Boltz-2 co-folder, Tensor-DTI remains competitive on CDK2 and improves the screening budget required to recover moderate fractions of high-affinity ligands on out-of-family targets under strict family-holdout splits. Additionally, we explore its applicability to protein-RNA and peptide-protein interactions. Our findings highlight the benefits of integrating multimodal information with contrastive objectives to enhance interaction-prediction accuracy and to provide more interpretable and reliability-aware models for virtual screening.

药物发现多模态学习对比学习生物互作预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。