arXiv:2601.09693cs.LGstat.ML2026-01

统一结构与配体数据的对比几何学习模型,提升药物设计效率。

Contrastive Geometric Learning Unlocks Unified Structure- and Ligand-Based Drug Design

  • 用对比几何学习统一蛋白与配体建模,无需预定义结合位点。
  • 零样本虚拟筛选表现优异,在靶点挖掘任务中显著超越现有方法。
  • 适合药物发现领域研究者,推动通用药物设计基础模型发展。

结构基与配体基计算药物设计传统上依赖分离的数据源和建模假设,限制了其大规模联合应用。本文提出对比几何学习统一药物设计(ConGLUDe),一种单一对比几何模型,可整合结构与配体训练。ConGLUDe结合几何蛋白编码器生成全蛋白表示及预测结合位点隐式嵌入,与快速配体编码器协同工作,无需预定义口袋。通过对比学习将配体与全局蛋白表示及多个候选结合位点对齐,支持配体条件下的口袋预测、虚拟筛选与靶点挖掘,且在蛋白-配体复合物与大规模生物活性数据上联合训练。在多样基准测试中,ConGLUDe实现具有竞争力的零样本虚拟筛选性能,在挑战性靶点挖掘任务中显著优于现有方法,并展现出最先进的配体条件口袋选择能力。结果凸显统一结构-配体训练的优势,为药物发现通用基础模型迈出关键一步。

原文摘要 · Abstract (English)

Structure-based and ligand-based computational drug design have traditionally relied on disjoint data sources and modeling assumptions, limiting their joint use at scale. In this work, we introduce Contrastive Geometric Learning for Unified Computational Drug Design (ConGLUDe), a single contrastive geometric model that unifies structure- and ligand-based training. ConGLUDe couples a geometric protein encoder that produces whole-protein representations and implicit embeddings of predicted binding sites with a fast ligand encoder, removing the need for predefined pockets. By aligning ligands with both global protein representations and multiple candidate binding sites through contrastive learning, ConGLUDe supports ligand-conditioned pocket prediction in addition to virtual screening and target fishing, while being trained jointly on protein-ligand complexes and large-scale bioactivity data. Across diverse benchmarks, ConGLUDe achieves competitive zero-shot virtual screening performance, substantially outperforms existing methods on a challenging target fishing task, and demonstrates state-of-the-art ligand-conditioned pocket selection. These results highlight the advantages of unified structure-ligand training and position ConGLUDe as a step toward general-purpose foundation models for drug discovery.

药物设计对比学习几何深度学习虚拟筛选

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。