用大模型和强化学习生成更有效的药物分子,成功率更高。
DrugGen: Advancing Drug Discovery with Large Language Models and Reinforcement Learning Feedback
- 基于DrugGPT改进,融合奖励反馈优化分子生成
- 100%生成有效结构,结合力预测值提升至7.22
- 适合药物研发、新药设计及老药新用研究者
传统药物设计受化学与生物复杂性制约,临床试验失败率高。深度学习中的生成模型如DrugGPT虽有潜力,但常生成无效结构且缺乏已批准药物特征,导致效率低下。为此,本文提出DrugGen,基于DrugGPT架构,通过在已批准药物-靶点互作数据上微调,并采用近端策略优化(PPO)结合预训练的蛋白-配体结合亲和力预测模型PLAPT及定制化无效结构评估器进行奖励反馈。多靶点评估显示,DrugGen实现100%有效结构生成(对比DrugGPT的95.5%),预测结合亲和力达7.22(6.30-8.07),优于DrugGPT的5.81(4.97-6.63),同时保持分子多样性与新颖性。对接模拟验证其能有效靶向结合位点,如对脂肪酸结合蛋白5(FABP5),生成分子得分分别为-9.537和-8.399,显著优于参考分子棕榈酸(-6.177)。此外,该模型还具备药物重定位与新型药效团设计潜力,为药物研发提供高效平台。
原文摘要 · Abstract (English)
Traditional drug design faces significant challenges due to inherent chemical and biological complexities, often resulting in high failure rates in clinical trials. Deep learning advancements, particularly generative models, offer potential solutions to these challenges. One promising algorithm is DrugGPT, a transformer-based model, that generates small molecules for input protein sequences. Although promising, it generates both chemically valid and invalid structures and does not incorporate the features of approved drugs, resulting in time-consuming and inefficient drug discovery. To address these issues, we introduce DrugGen, an enhanced model based on the DrugGPT structure. DrugGen is fine-tuned on approved drug-target interactions and optimized with proximal policy optimization. By giving reward feedback from protein-ligand binding affinity prediction using pre-trained transformers (PLAPT) and a customized invalid structure assessor, DrugGen significantly improves performance. Evaluation across multiple targets demonstrated that DrugGen achieves 100% valid structure generation compared to 95.5% with DrugGPT and produced molecules with higher predicted binding affinities (7.22 [6.30-8.07]) compared to DrugGPT (5.81 [4.97-6.63]) while maintaining diversity and novelty. Docking simulations further validate its ability to generate molecules targeting binding sites effectively. For example, in the case of fatty acid-binding protein 5 (FABP5), DrugGen generated molecules with superior docking scores (FABP5/11, -9.537 and FABP5/5, -8.399) compared to the reference molecule (Palmitic acid, -6.177). Beyond lead compound generation, DrugGen also shows potential for drug repositioning and creating novel pharmacophores for existing targets. By producing high-quality small molecules, DrugGen provides a high-performance medium for advancing pharmaceutical research and drug discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。