用新扩散模型生成抗体序列,减少对天然基因序列的依赖,提升设计质量。
Conditional generation of antibody sequences with classifier-guided germline-absorbing discrete diffusion

- 引入胚系吸收扩散机制,让模型只学习从胚系到成熟序列的生物学路径。
- 非胚系残基预测准确率从26%提升至46%,接近生物真实变异上限。
- 支持任意分类器引导生成,适合优化亲和力和疏水性等属性的设计任务。
抗体药物是现代最成功的疗法之一,但计算设计兼具理想结合能力与可开发性的抗体仍具挑战。尽管蛋白质语言模型(pLMs)已成为抗体序列设计的强大工具,现有方法普遍存在两大局限:主要记忆胚系序列而非建模有意义的体细胞变异,且难以灵活支持分类器引导的条件生成。本文提出两项关键贡献:首先,离散扩散微调在抗体序列上实现强语言建模性能,并支持任意现成分类器的条件生成;其次,提出胚系吸收扩散,一种新型离散扩散噪声过程,以胚系序列而非掩码序列作为吸收态。这一生物学启发的归纳偏置使模型仅学习从胚系到观测序列的轨迹,有效排除遗传变异与V(D)J重排统计的影响,显著缓解胚系偏差。实验显示,该方法将非胚系残基预测准确率从26%提升至46%,接近真实生物变异的理论上限。进一步在改善疏水性和预测结合亲和力的条件生成任务中验证,模型在分类器遵循度与样本质量间取得更优权衡,显著优于EvoProtGrad等主流梯度驱动采样策略。
原文摘要 · Abstract (English)
Antibody therapeutics are among the most successful modern medicines, yet computationally designing antibodies with desirable binding and developability properties remains challenging. While protein language models (pLMs) have emerged as powerful tools for antibody sequence design, existing approaches largely suffer from two key limitations: they predominantly memorize germline sequences rather than modeling biologically meaningful somatic variation, and they offer limited support for flexible classifier-guided conditional generation. We address these challenges through two primary contributions. First, we demonstrate that discrete diffusion fine-tuning achieves strong language modeling performance on antibody sequences while allowing for generation conditioned on any off-the-shelf classifier. Second, we introduce germline absorbing diffusion, a novel modification of the discrete diffusion noise process in which the germline sequence - rather than a masked sequence - serves as the absorbing state. This biologically motivated inductive bias restricts the model to learning the trajectory from germline to observed sequence, effectively excluding genetic variation and V(D)J recombination statistics from the learned distribution and dramatically mitigating germline bias. We show that germline diffusion improves non-germline residue prediction accuracy from 26 percent to 46 percent, approaching the theoretical upper bound set by true biological variability. We then demonstrate the utility of our germline diffusion model on the conditional generation tasks of sampling antibodies with improved hydrophobicity and predicted binding affinity. On both tasks our model shows an improved tradeoff between class adherence and sample quality, significantly outperforming EvoProtGrad, a popular strategy to sample from pLMs with gradient-based discrete Markov Chain Monte Carlo.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。