arXiv:2506.15309cs.LGcs.AI2025-06被引 1

用主动学习优化多靶点抑制剂生成,提升多样性与有效性

Active Learning-Guided Seq2Seq Variational Autoencoder for Multi-target Inhibitor Generation

  • 结合序列到序列变分自编码器与主动学习,迭代扩展化学空间
  • 在三个冠状病毒主蛋白酶上生成结构多样且具广谱抑制活性的候选分子
  • 适合药物设计中需兼顾多靶点、低奖励信号的复杂场景

同时优化分子对多个治疗靶点的活性仍是药物发现中的重大挑战,尤其因奖励稀疏和设计约束冲突。我们提出一种结构化的主动学习范式,将序列到序列变分自编码器(Seq2Seq VAE)嵌入迭代循环,以平衡化学多样性、分子质量与多靶点亲和力。该方法交替扩展潜在空间中的化学可行区域,并基于日益严格的多靶点对接阈值逐步约束分子。在针对三种相关冠状病毒主蛋白酶(SARS-CoV-2、SARS-CoV、MERS-CoV)的验证研究中,该方法高效生成了结构多样化的泛抑制剂候选分子。研究表明,在主动学习流程中合理安排化学过滤时机与位置,显著增强了有益化学空间的探索能力,使稀疏奖励、多目标的药物设计问题变为可计算处理的任务。本框架为高效导航复杂多药理学景观提供了通用路线。

原文摘要 · Abstract (English)

Simultaneously optimizing molecules against multiple therapeutic targets remains a profound challenge in drug discovery, particularly due to sparse rewards and conflicting design constraints. We propose a structured active learning (AL) paradigm integrating a sequence-to-sequence (Seq2Seq) variational autoencoder (VAE) into iterative loops designed to balance chemical diversity, molecular quality, and multi-target affinity. Our method alternates between expanding chemically feasible regions of latent space and progressively constraining molecules based on increasingly stringent multi-target docking thresholds. In a proof-of-concept study targeting three related coronavirus main proteases (SARS-CoV-2, SARS-CoV, MERS-CoV), our approach efficiently generated a structurally diverse set of pan-inhibitor candidates. We demonstrate that careful timing and strategic placement of chemical filters within this active learning pipeline markedly enhance exploration of beneficial chemical space, transforming the sparse-reward, multi-objective drug design problem into an accessible computational task. Our framework thus provides a generalizable roadmap for efficiently navigating complex polypharmacological landscapes.

药物生成主动学习多靶点VAE

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。