用序列和结构联合学习,提升药物筛选准确率
S$^2$Drug: Bridging Protein Sequence and 3D Structure in Contrastive Representation Learning for Virtual Screening
- 分两阶段:先用化学数据预训练序列,再融合结构信息微调
- 在多个基准上优于现有方法,结合序列与结构使匹配更精准
- 适合做药物发现中蛋白-配体对接的算法研究者
虚拟筛选是药物发现中的关键任务,旨在识别能与特定蛋白口袋结合的小分子配体。现有深度学习方法多依赖结构数据而忽略更易获取的蛋白序列,导致泛化能力受限。大规模蛋白-配体数据集中的冗余和噪声使直接整合序列信息面临挑战。为此,我们提出S²Drug,一种两阶段框架,显式融合蛋白序列与三维结构上下文,在对比表示学习中实现蛋白-配体匹配。第一阶段在ChemBL上基于ESM2骨干网络进行序列预训练,采用定制采样策略降低蛋白与配体侧的冗余和噪声。第二阶段在PDBBind上微调,通过残基级门控模块融合序列与结构信息,并引入辅助结合位点预测任务。该任务引导模型精确定位蛋白序列中的结合残基及其三维空间排列,从而优化蛋白-配体匹配。在多个基准测试中,S²Drug持续提升虚拟筛选性能,并在结合位点预测上取得优异结果,验证了序列与结构融合在对比学习中的价值。
原文摘要 · Abstract (English)
Virtual screening (VS) is an essential task in drug discovery, focusing on the identification of small-molecule ligands that bind to specific protein pockets. Existing deep learning methods, from early regression models to recent contrastive learning approaches, primarily rely on structural data while overlooking protein sequences, which are more accessible and can enhance generalizability. However, directly integrating protein sequences poses challenges due to the redundancy and noise in large-scale protein-ligand datasets. To address these limitations, we propose \textbf{S$^2$Drug}, a two-stage framework that explicitly incorporates protein \textbf{S}equence information and 3D \textbf{S}tructure context in protein-ligand contrastive representation learning. In the first stage, we perform protein sequence pretraining on ChemBL using an ESM2-based backbone, combined with a tailored data sampling strategy to reduce redundancy and noise on both protein and ligand sides. In the second stage, we fine-tune on PDBBind by fusing sequence and structure information through a residue-level gating module, while introducing an auxiliary binding site prediction task. This auxiliary task guides the model to accurately localize binding residues within the protein sequence and capture their 3D spatial arrangement, thereby refining protein-ligand matching. Across multiple benchmarks, S$^2$Drug consistently improves virtual screening performance and achieves strong results on binding site prediction, demonstrating the value of bridging sequence and structure in contrastive learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。