arXiv:2509.00684cs.LGcs.AI2025-09

用对比学习引导分子生成,低数据下高效设计新药分子。

Valid Property-Enhanced Contrastive Learning for Targeted Optimization & Resampling for Novel Drug Design

  • 通过属性增强对比学习,让生成模型聚焦有效化学空间。
  • 生成8374个分子,100个超越-15.0 kcal/mol阈值,最优达-17.6 kcal/mol。
  • 适合药物发现中数据少、需可解释设计的科研人员使用。

在低数据条件下,如何高效引导生成模型聚焦药理相关化学空间仍是分子药物发现的核心挑战。我们提出VECTOR+:一种结合属性引导表示学习与可控分子生成的框架,适用于回归与分类任务,支持可解释、数据高效的化学功能空间探索。在两个数据集上评估:一个含296个化合物的PD-L1抑制剂集(有实验IC50值),另一个按结合模式划分的激酶抑制剂集(2,056分子)。尽管训练数据有限,VECTOR+仍生成新颖且可合成的候选分子。针对PD-L1(PDB 5J89),8,374个生成分子中有100个超过-15.0 kcal/mol打分阈值,最佳得分-17.6 kcal/mol,优于最优参考抑制剂(-15.4 kcal/mol)。最优分子保留关键双苯基药效团并引入新结构。250纳秒分子动力学模拟显示结合稳定(配体RMSD < 2.5 Å)。VECTOR+在激酶抑制剂上也表现优异,生成分子打分高于布格替尼和索拉非尼等现有药物。相较于JT-VAE和MolGPT,在打分、新颖性、唯一性和Tanimoto相似度上均更优。本工作为低数据环境下属性约束的分子设计提供了稳健、可扩展的方案,实现了对比学习与生成建模的融合,推动可复现的AI加速药物发现。

原文摘要 · Abstract (English)

Efficiently steering generative models toward pharmacologically relevant regions of chemical space remains a major obstacle in molecular drug discovery under low-data regimes. We present VECTOR+: Valid-property-Enhanced Contrastive Learning for Targeted Optimization and Resampling, a framework that couples property-guided representation learning with controllable molecule generation. VECTOR+ applies to both regression and classification tasks and enables interpretable, data-efficient exploration of functional chemical space. We evaluate on two datasets: a curated PD-L1 inhibitor set (296 compounds with experimental $IC_{50}$ values) and a receptor kinase inhibitor set (2,056 molecules by binding mode). Despite limited training data, VECTOR+ generates novel, synthetically tractable candidates. Against PD-L1 (PDB 5J89), 100 of 8,374 generated molecules surpass a docking threshold of $-15.0$ kcal/mol, with the best scoring $-17.6$ kcal/mol compared to the top reference inhibitor ($-15.4$ kcal/mol). The best-performing molecules retain the conserved biphenyl pharmacophore while introducing novel motifs. Molecular dynamics (250 ns) confirm binding stability (ligand RMSD < $2.5$ angstroms). VECTOR+ generalizes to kinase inhibitors, producing compounds with stronger docking scores than established drugs such as brigatinib and sorafenib. Benchmarking against JT-VAE and MolGPT across docking, novelty, uniqueness, and Tanimoto similarity highlights the superior performance of our method. These results position our work as a robust, extensible approach for property-conditioned molecular design in low-data settings, bridging contrastive learning and generative modeling for reproducible, AI-accelerated discovery.

分子生成对比学习药物设计低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。