用对比学习生成可合成药物分子,提升靶点结合预测精度
Structure-guided molecular design with contrastive 3D protein-ligand learning

- 通过等变Transformer对比学习编码蛋白-配体三维结构
- 生成分子在多个靶点上预测结合性能优异且可合成
- 适合需要高效设计新药的科研人员使用
基于结构的药物发现面临两大挑战:准确捕捉三维蛋白-配体相互作用,并在超大规模化学空间中筛选出可合成的候选分子。本文提出统一框架,结合对比3D结构编码与基于商业化合物空间的自回归分子生成。首先,引入SE(3)-等变Transformer,通过对比学习将配体与口袋结构编码至共享嵌入空间,在零样本虚拟筛选中表现优异。其次,将这些嵌入融入多模态化学语言模型(MCLM),根据口袋或配体结构生成目标特异性分子,通过学习到的数据集标记引导输出至目标化学空间,生成候选分子在多个靶点上均展现出良好的预测结合性能。
原文摘要 · Abstract (English)
Structure-based drug discovery faces the dual challenge of accurately capturing 3D protein-ligand interactions while navigating ultra-large chemical spaces to identify synthetically accessible candidates. In this work, we present a unified framework that addresses these challenges by combining contrastive 3D structure encoding with autoregressive molecular generation conditioned on commercial compound spaces. First, we introduce an SE(3)-equivariant transformer that encodes ligand and pocket structures into a shared embedding space via contrastive learning, achieving competitive results in zero-shot virtual screening. Second, we integrate these embeddings into a multimodal Chemical Language Model (MCLM). The model generates target-specific molecules conditioned on either pocket or ligand structures, with a learned dataset token that steers the output toward targeted chemical spaces, yielding candidates with favorable predicted binding properties across diverse targets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。