对比不同模型架构对抗体特征的捕捉能力,发现专用模型更懂抗体关键区域。
Exploring Protein Language Model Architecture-Induced Biases for Antibody Comprehension
- 用注意力分析揭示抗体模型天然聚焦抗原结合区
- 专用模型在抗体特异性预测中表现一致但偏见各异
- 适合抗体设计与药物研发领域的研究人员参考
蛋白质语言模型(PLMs)在理解蛋白序列方面取得显著进展,但不同模型架构对抗体特异性生物学特征的捕捉程度尚不明确。本文系统评估了三种先进PLM(AntiBERTa、BioBERT、ESM2)与通用语言模型(GPT-2)在抗体靶点特异性预测任务上的表现。结果表明,尽管所有模型均达到高分类准确率,但在识别V基因使用、体细胞超突变模式和同种型信息等生物特征时表现出明显偏差。通过注意力归因分析发现,如AntiBERTa等抗体专用模型能自然聚焦于互补决定区(CDRs),而通用蛋白模型则需显式引入以CDR为中心的训练策略才能有效学习。研究揭示了模型架构与生物特征提取之间的关联,为计算抗体设计中的未来PLM开发提供重要指导。
原文摘要 · Abstract (English)
Recent advances in protein language models (PLMs) have demonstrated remarkable capabilities in understanding protein sequences. However, the extent to which different model architectures capture antibody-specific biological properties remains unexplored. In this work, we systematically investigate how architectural choices in PLMs influence their ability to comprehend antibody sequence characteristics and functions. We evaluate three state-of-the-art PLMs-AntiBERTa, BioBERT, and ESM2--against a general-purpose language model (GPT-2) baseline on antibody target specificity prediction tasks. Our results demonstrate that while all PLMs achieve high classification accuracy, they exhibit distinct biases in capturing biological features such as V gene usage, somatic hypermutation patterns, and isotype information. Through attention attribution analysis, we show that antibody-specific models like AntiBERTa naturally learn to focus on complementarity-determining regions (CDRs), while general protein models benefit significantly from explicit CDR-focused training strategies. These findings provide insights into the relationship between model architecture and biological feature extraction, offering valuable guidance for future PLM development in computational antibody design.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。