提升抗体-抗原复合物建模精度,解决现有方法在结合界面预测上的短板。
Computational Modeling of Antibody-Antigen Complexes: PLM-Based and MSA-Based Approaches

- 基于蛋白质语言模型的嵌入特征,提升抗体单体结构预测准确率
- 通过MSA优化与收敛感知重采样,显著增强复合物预测性能
- 无需重新训练模型,适用于现有工具如AlphaFold3的改进
抗体通过特异性识别并中和抗原,在免疫反应中起核心作用,治疗性抗体已成为癌症和自身免疫疾病的主要药物。然而其发现仍依赖大量体外筛选,精准的计算建模可优先候选分子、减轻实验负担并加速理性设计。尽管蛋白质及复合物预测取得进展,抗体相关任务仍存在明显性能差距,制约下游应用。本研究探究抗体任务更难的原因,并从两条互补路径提出改进:首先,采用多蛋白质语言模型(PLM)嵌入,在抗体单体预测中达到当前最优的CDR-H3准确率;但扩展至复合物预测时表现不佳,因缺乏抗体与抗原间的共进化信号,单序列PLM表示难以可靠识别结合界面。其次,提出两种基于多序列比对(MSA)的干预策略:MSA精炼(聚焦CDR区域过滤并从更大数据库恢复深度)与收敛感知重采样(选择稳定中间状态用于最终扩散采样),两者联合在独立测试集上持续优于AlphaFold3基线。由于仅修改MSA构建与重采样行为,不涉及模型参数,因此无需重新训练或权重访问即可应用。
原文摘要 · Abstract (English)
Antibodies play a central role in the immune response by specifically recognizing and neutralizing antigens, and therapeutic antibodies have become major drugs for cancer and autoimmune diseases. However, their discovery still relies on extensive in vitro screening, and accurate computational modeling of antibody structures and antibody-antigen interactions can prioritize candidates, reduce experimental burden, and accelerate rational design. Despite recent advances in high-accuracy protein and complex prediction, a persistent performance gap remains for antibody-related tasks compared with general protein-protein interactions, limiting downstream design. This thesis investigates why antibody-related tasks are harder and proposes improvements along two complementary directions. First, we investigate protein language model (PLM)-based methods for antibody and antibody-antigen structure prediction. Using embeddings from multiple PLMs, our approach achieves the best CDR-H3 accuracy among compared PLM-based methods on antibody monomer prediction. Extending it to complex prediction does not generalize: without co-evolutionary signals between antibody and antigen, single-sequence PLM representations do not reliably identify binding interfaces. Second, we develop two MSA-based interventions for antibody-antigen complex prediction: MSA refinement, which combines CDR-focused filtering with depth recovery from a larger sequence database, and convergence-aware recycling, which selects a stable intermediate recycle state for final diffusion sampling. Together, these interventions provide consistent gains over the AlphaFold3 baseline on a held-out antibody-antigen test set. Because the methods modify MSA construction and recycling behavior rather than model parameters, they apply without retraining or weight access.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。