用生成模型直接产出有活性的候选分子,验证其替代传统筛选的可行性。
From In Silico to In Vitro: Evaluating Molecule Generative Models for Hit Generation
- 构建多阶段筛选框架,综合物理化学、结构与生物活性标准定义类先导化合物空间。
- 三种生成模型在多个靶点上生成有效且多样、具有生物活性的分子,部分已体外验证。
- 首次将类先导分子生成作为独立任务评估,适合药物发现与生成模型研究者。
先导化合物识别是药物发现中关键但成本高昂的环节,传统依赖大规模化合物库的高通量筛选。尽管虚拟筛选有所进展,仍耗时且昂贵。深度学习推动了生成模型的发展,使其能学习复杂分子表示并从头生成新化合物。然而,用机器学习完全替代整个药物发现流程仍具挑战。本文聚焦于是否可用生成模型替代其中一环——类先导分子生成。据我们所知,这是首个将类先导分子生成明确视为独立任务,并实证检验生成模型能否直接支持或取代传统先导识别流程的研究。我们提出一个针对性评估框架,结合理化性质、结构特征和生物活性标准,构建多阶段过滤管道以界定类先导化学空间。比较了两种自回归模型与一种基于扩散的生成模型,在多种数据集与训练设置下的表现,使用标准指标及靶点特异性对接评分评估输出结果。结果显示,这些模型能在多个靶点上生成有效、多样且具有生物相关性的化合物,少数针对GSK-3β的命中分子已成功合成并在体外确认活性。同时,我们指出当前评估指标与训练数据的关键局限。
原文摘要 · Abstract (English)
Hit identification is a critical yet resource-intensive step in the drug discovery pipeline, traditionally relying on high-throughput screening of large compound libraries. Despite advancements in virtual screening, these methods remain time-consuming and costly. Recent progress in deep learning has enabled the development of generative models capable of learning complex molecular representations and generating novel compounds de novo. However, using ML to replace the entire drug-discovery pipeline is highly challenging. In this work, we rather investigate whether generative models can replace one step of the pipeline: hit-like molecule generation. To the best of our knowledge, this is the first study to explicitly frame hit-like molecule generation as a standalone task and empirically test whether generative models can directly support this stage of the drug discovery pipeline. Specifically, we investigate if such models can be trained to generate hit-like molecules, enabling direct incorporation into, or even substitution of, traditional hit identification workflows. We propose an evaluation framework tailored to this task, integrating physicochemical, structural, and bioactivity-related criteria within a multi-stage filtering pipeline that defines the hit-like chemical space. Two autoregressive and one diffusion-based generative models were benchmarked across various datasets and training settings, with outputs assessed using standard metrics and target-specific docking scores. Our results show that these models can generate valid, diverse, and biologically relevant compounds across multiple targets, with a few selected GSK-3$β$ hits synthesized and confirmed active in vitro. We also identify key limitations in current evaluation metrics and available training data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。