用扩散模型生成环肽,兼顾化学合理性与靶点结合能力
PepALD: Macrocyclic Peptide Generation via Autoregressive Latent Diffusion

- 基于结构化化学嵌入的自回归扩散框架,逐残基生成环肽
- 在虚拟实验中生成质量优于主流基线模型,且能优化结合亲和力
- 适合药物设计人员探索新型环肽分子,尤其关注膜渗透性
环状肽是针对胞内靶点的有前途治疗候选物,但其设计需同时控制非天然单体化学、环拓扑结构、膜渗透性和靶点结合能力。现有基于SMILES或HELM字符串的生成模型要么在长原子序列空间中操作,要么将单体视为缺乏化学基础的符号标记。我们提出PepALD,一种用于从头生成环状肽的自回归潜空间扩散(ALD)基础模型。该模型以结构化化学嵌入表示HELM单体,通过化学感知潜空间中的上下文条件扩散生成每个残基,在自回归生成过程中预测考虑R基团的环闭合,并使用赢家保护的扩散适应偏好优化对去噪器进行亲和力奖励对齐。体外实验表明,PepALD在生成质量和奖励优化性能方面均优于代表性肽生成基线。
原文摘要 · Abstract (English)
Macrocyclic peptides are promising therapeutic candidates for intracellular targets, but their design requires simultaneous control over non-natural monomer chemistry, ring topology, membrane permeability, and target binding. Existing SMILES- or HELM-string generative models either operate in long atom-level sequence spaces or treat monomers as symbolic tokens with limited chemical grounding. We introduce PepALD, an Autoregressive Latent Diffusion (ALD) foundation model for \textit{de novo} macrocyclic peptide generation. The model represents HELM monomers with structured chemical embeddings, generates each residue through context-conditioned diffusion in chemically informed latent space, predicts R-group-aware ring closures during autoregressive generation, and aligns the denoiser to affinity rewards using winner-protected diffusion-adapted preference optimization. In silico experiments demonstrate PepALD's generation quality and reward-optimization performance against representative peptide generation baselines.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。