用进化与结构先验结合,解决抗体互补位设计中氨基酸种类丢失问题。
EvoStruct: Bridging Evolutionary and Structural Priors for Antibody CDR Design via Protein Language Model Adaptation

- 通过交叉注意力适配器融合语言模型与三维结构信息
- 在CHIMERA-Bench上序列恢复率提升16%,困惑度降低43%
- 适合抗体设计、药物研发人员关注,尤其需高多样性序列的场景
针对抗体互补决定区(CDR)设计中,等变图神经网络(GNN)方法虽能实现最高序列恢复率,但存在严重词汇坍缩问题。当前最优的GNN方法过度预测少数氨基酸(如酪氨酸和甘氨酸),忽略功能关键残基。我们发现根源在于GNN编码器仅从有限结构数据中重新学习氨基酸分布,忽略了进化数据库中的替代模式。为此,提出EvoStruct,通过跨注意力适配器将冻结的蛋白质语言模型(PLM)与E(3)-等变GNN提供的三维结构上下文相结合。不同于通用蛋白设计的PLM-结构适配器,EvoStruct通过渐进式解冻PLM和R-Drop一致性正则化,专门解决CDR设计中的词汇坍缩问题。在CHIMERA-Bench数据集上,EvoStruct在氨基酸恢复率和困惑度方面均优于多种抗体设计方法,相比最佳基线,序列恢复率提升16%,困惑度降低43%,氨基酸多样性提高2.3倍,且与真实结合对的相关性最高。
原文摘要 · Abstract (English)
Equivariant graph neural network (GNN) methods for antibody complementarity-determining region (CDR) design achieve the highest sequence recovery but suffer from severe vocabulary collapse. The current best GNN methods over-predict very few amino acids, such as tyrosine and glycine, while ignoring functionally important residues. We trace this failure to GNN encoders learning amino acid distributions de novo from limited structural data, discarding substitution patterns encoded in evolutionary databases. To resolve this, we propose EvoStruct, which bridges a frozen protein language model (PLM) with 3D structural context from an E(3)-equivariant GNN via a cross-attention adapter. Unlike prior PLM-structure adapters for general protein design, EvoStruct targets the vocabulary collapse problem specific to CDR design through progressive PLM unfreezing and R-Drop consistency regularization. On the CHIMERA-Bench dataset, EvoStruct achieves the highest amino acid recovery and lowest perplexity among several antibody design methods, improving sequence recovery by 16% and reducing perplexity by 43% relative to the best GNN baselines, while recovering 2.3x greater amino acid diversity and the highest binding-pair correlation with ground truth.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。