arXiv:2604.20019cs.LG2026-04

用强化学习生成共价抑制剂,同时优化结合力与选择性。

Multi-Objective Reinforcement Learning for Generating Covalent Inhibitor Candidates

论文配图:Multi-Objective Reinforcement Learning for Generating Covalent Inhibitor Candidates
图 1 · 摘自论文原文
  • 基于SMILES的LSTM模型,通过多目标强化学习生成分子。
  • 在10,000个结构中复现已知抑制剂率最高达0.74%。
  • 自发生成新反应基团,适合药物研发人员使用。

合理设计共价抑制剂需同时优化结合亲和力、靶点选择性及亲电活性等多重性质,传统筛选难以应对。本文提出一种基于多目标强化学习(RL)的生成管道,应用于表皮生长因子受体(EGFR)和乙酰胆碱酯酶(ACHE)两个靶点。采用基于SMILES的预训练LSTM作为生成模型,通过策略梯度优化,结合帕累托拥挤距离平衡合成可及性、预测共价活性、残基亲和力及近似对接评分等多重目标。在10,000结构的生成中,对EGFR和ACHE的已知抑制剂复现率分别达0.50%和0.74%;经进一步对接筛选,候选分子与靶点残基距离最短分别为5.5 Å(EGFR)和3.2 Å(ACHE)。更显著的是,该方法自发生成了训练数据中未包含的反应基团,如烯烃、3-氧代-β-磺内酰胺和α-亚甲基-β-内酯,均在文献中有共价反应基团支持。结果表明,该方法可探索超越训练分布的共价化学空间,对共价药物发现具有实用价值。

原文摘要 · Abstract (English)

Rational design of covalent inhibitors requires simultaneously optimizing multiple properties, such as binding affinity, target selectivity, or electrophilic reactivity. This presents a multi-objective problem not easily addressed by screening alone. Here we present a machine learning pipeline for generating covalent inhibitor candidates using multi-objective reinforcement learning (RL), applied to two targets: epidermal growth factor receptor (EGFR) and acetylcholinesterase (ACHE). A SMILES-based pretrained LSTM serves as the generative model, optimized via policy gradient RL with Pareto crowding distance to balance competing scoring functions including synthetic accessibility, predicted covalent activity, residue affinity, and an approximated docking score. The pipeline rediscovers known covalent inhibitors at rates of up to 0.50% (EGFR) and 0.74% (ACHE) in 10,000-structure runs, with candidate structures achieving warhead-to-residue distances as short as 5.5 angstrom (EGFR) and 3.2 angstrom (ACHE) after further docking-based screening. More notably, the pipeline spontaneously generates structures bearing warhead motifs absent from the training data - including allenes, 3-oxo-$β$-sultams, and $α$-methylene-$β$-lactones - all of which have independent literature support as covalent warheads. These results suggest that RL-guided generation can explore covalent chemical space beyond its training distribution, and may be useful as a tool for medicinal chemists working on covalent drug discovery.

药物设计强化学习共价抑制剂生成模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。