用大模型+推理+强化学习,让环肽设计更智能可解释
PepThink-R1: LLM for Interpretable Cyclic Peptide Optimization with CoT SFT and Reinforcement Learning
- 基于思维链和强化学习,逐个优化氨基酸组合
- 生成的环肽脂溶性、稳定性、暴露量显著提升
- 适合药物研发人员快速设计高潜力肽类分子
设计具有特定性质的治疗性肽类受制于序列空间庞大、实验数据有限以及现有生成模型可解释性差。为此,我们提出PepThink-R1,一种将大语言模型(LLM)与思维链(CoT)监督微调及强化学习(RL)结合的生成框架。不同于以往方法,PepThink-R1在序列生成过程中显式推理单体级修饰,实现可解释的设计决策,同时优化多种药理性质。通过定制奖励函数平衡化学有效性与性质提升,模型自主探索多样序列变体。实验表明,PepThink-R1生成的环肽在脂溶性、稳定性和暴露量上显著优于现有通用大模型(如GPT-5)和领域专用基线模型,在优化成功率和可解释性方面均表现更优。据我们所知,这是首个结合显式推理与强化学习驱动性质控制的基于大模型的肽设计框架,为治疗发现中的可靠、透明肽类优化迈出关键一步。
原文摘要 · Abstract (English)
Designing therapeutic peptides with tailored properties is hindered by the vastness of sequence space, limited experimental data, and poor interpretability of current generative models. To address these challenges, we introduce PepThink-R1, a generative framework that integrates large language models (LLMs) with chain-of-thought (CoT) supervised fine-tuning and reinforcement learning (RL). Unlike prior approaches, PepThink-R1 explicitly reasons about monomer-level modifications during sequence generation, enabling interpretable design choices while optimizing for multiple pharmacological properties. Guided by a tailored reward function balancing chemical validity and property improvements, the model autonomously explores diverse sequence variants. We demonstrate that PepThink-R1 generates cyclic peptides with significantly enhanced lipophilicity, stability, and exposure, outperforming existing general LLMs (e.g., GPT-5) and domain-specific baseline in both optimization success and interpretability. To our knowledge, this is the first LLM-based peptide design framework that combines explicit reasoning with RL-driven property control, marking a step toward reliable and transparent peptide optimization for therapeutic discovery.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。