arXiv:2604.19562cs.LG2026-04

用对比学习生成可合成药物分子,提升靶点结合预测精度

Structure-guided molecular design with contrastive 3D protein-ligand learning

论文配图:Structure-guided molecular design with contrastive 3D protein-ligand learning
图 1 · 摘自论文原文
  • 通过等变Transformer对比学习编码蛋白-配体三维结构
  • 生成分子在多个靶点上预测结合性能优异且可合成
  • 适合需要高效设计新药的科研人员使用

基于结构的药物发现面临两大挑战:准确捕捉三维蛋白-配体相互作用,并在超大规模化学空间中筛选出可合成的候选分子。本文提出统一框架,结合对比3D结构编码与基于商业化合物空间的自回归分子生成。首先,引入SE(3)-等变Transformer,通过对比学习将配体与口袋结构编码至共享嵌入空间,在零样本虚拟筛选中表现优异。其次,将这些嵌入融入多模态化学语言模型(MCLM),根据口袋或配体结构生成目标特异性分子,通过学习到的数据集标记引导输出至目标化学空间,生成候选分子在多个靶点上均展现出良好的预测结合性能。

原文摘要 · Abstract (English)

Structure-based drug discovery faces the dual challenge of accurately capturing 3D protein-ligand interactions while navigating ultra-large chemical spaces to identify synthetically accessible candidates. In this work, we present a unified framework that addresses these challenges by combining contrastive 3D structure encoding with autoregressive molecular generation conditioned on commercial compound spaces. First, we introduce an SE(3)-equivariant transformer that encodes ligand and pocket structures into a shared embedding space via contrastive learning, achieving competitive results in zero-shot virtual screening. Second, we integrate these embeddings into a multimodal Chemical Language Model (MCLM). The model generates target-specific molecules conditioned on either pocket or ligand structures, with a learned dataset token that steers the output toward targeted chemical spaces, yielding candidates with favorable predicted binding properties across diverse targets.

分子生成蛋白质-配体对比学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。