arXiv:2604.24474cs.LG2026-04中稿 · ICML被引 1

用预训练分子嵌入距离提升药物筛选与生成效率

Advancing Ligand-based Virtual Screening and Molecular Generation with Pretrained Molecular Embedding Distance

论文配图:Advancing Ligand-based Virtual Screening and Molecular Generation with Pretrained Molecular Embedding Distance
图 1 · 摘自论文原文
  • 直接使用预训练模型输出的嵌入距离,无需任务微调
  • 在虚拟筛选中排序效果优于传统方法,生成质量更高
  • 适合大规模药物发现场景,通用性强

分子相似性在基于配体的药物发现中至关重要,如虚拟筛选、类似物搜索和目标导向的分子生成。然而,传统相似性度量方法(从指纹的Tanimoto系数到3D形状叠加)往往在大规模应用时计算成本高,或依赖手工设计的分子描述符。许多深度学习方法仍需特定相似性监督或昂贵的数据标注,限制了其跨靶标的通用性。本文提出预训练嵌入距离(PED),可直接从预训练分子模型中计算,无需任务特定训练。实验表明,PED与传统相似性度量具有显著相关性,在虚拟筛选分子排序和通过奖励设计引导分子生成方面均表现优异。结果表明,预训练分子嵌入蕴含丰富的结构信息,可作为现代AI辅助药物发现中一种有前景且可扩展的相似性度量方式。

原文摘要 · Abstract (English)

Molecular similarity plays a central role in ligand-based drug discovery, such as virtual screening, analog searching, and goal-directed molecular generation. However, traditional similarity measures, ranging from fingerprint-based Tanimoto coefficients to 3D shape overlays, are often computationally expensive at scale or rely on hand-crafted molecular descriptors. Meanwhile, many deep learning approaches to similarity-aware design still depend on similarity-specific supervision or costly data curation, limiting their generality across targets. In this work, we propose pretrained embedding distance (PED) as an effective alternative, computed directly from pretrained molecular models without task-specific training. Experimental results show that PED exhibits distinct correlations with traditional similarity metrics, and performs effectively in both ranking molecules for virtual screening and guiding molecular generation via reward design. These findings suggest that pretrained molecular embeddings capture rich structural information and can serve as a promising and scalable similarity measurement for modern AI-aided drug discovery.

分子生成虚拟筛选预训练模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。