用强化学习优化生成抗菌肽,提升活性与安全性。
ProDCARL: Reinforcement Learning-Aligned Diffusion Models for De Novo Antimicrobial Peptide Design
- 结合扩散模型与奖励机制,定向生成抗菌肽。
- 预测活性提升至0.178,高质候选率达6.3%。
- 适合药物研发者快速筛选潜在抗菌分子。
抗菌耐药性威胁医疗可持续性,推动低成本计算发现抗菌肽(AMPs)。从头设计需兼顾抗菌活性与低毒性,但传统概率训练模型无法显式约束此目标。本文提出ProDCARL框架,将基于扩散的蛋白质生成器(EvoDiff OA-DM 38M)与AMP活性及毒性预测器结合,通过AMP序列微调扩散先验以获得领域感知生成器。采用top-k策略梯度更新,利用分类器奖励+熵正则化+早停机制,保持多样性并减少奖励欺骗。体外实验显示,生成物平均预测AMP得分由微调后的0.081提升至0.178;联合高质量命中率达6.3%,满足pAMP>0.7且pTox<0.3。生成序列均一性为1-均值成对身份0.929,保持高多样性。定性分析表明候选分子具备类AMP的结构与语义特征。ProDCARL可作为候选生成器缩小实验搜索空间,实验验证待后续开展。
原文摘要 · Abstract (English)
Antimicrobial resistance threatens healthcare sustainability and motivates low-cost computational discovery of antimicrobial peptides (AMPs). De novo peptide generation must optimize antimicrobial activity and safety through low predicted toxicity, but likelihood-trained generators do not enforce these goals explicitly. We introduce ProDCARL, a reinforcement-learning alignment framework that couples a diffusion-based protein generator (EvoDiff OA-DM 38M) with sequence property predictors for AMP activity and peptide toxicity. We fine-tune the diffusion prior on AMP sequences to obtain a domain-aware generator. Top-k policy-gradient updates use classifier-derived rewards plus entropy regularization and early stopping to preserve diversity and reduce reward hacking. In silico experiments show ProDCARL increases the mean predicted AMP score from 0.081 after fine-tuning to 0.178. The joint high-quality hit rate reaches 6.3\% with pAMP $>$0.7 and pTox $<$0.3. ProDCARL maintains high diversity, with $1-$mean pairwise identity equal to 0.929. Qualitative analyses with AlphaFold3 and ProtBERT embeddings suggest candidates show plausible AMP-like structural and semantic characteristics. ProDCARL serves as a candidate generator that narrows experimental search space, and experimental validation remains future work.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。