基于RoBERTa的抗体专用大模型,支持抗体设计与功能预测。
Antibody Foundational Model : Ab-RoBERTa
- 用RoBERTa架构构建抗体专属语言模型,参数量仅125M
- 在抗体序列数据集OAS上训练,支持表位预测与人源性评估
- 开源可用,适合抗体工程、药物研发人员快速部署
随着抗体类药物日益重要,抗体工程成为关键研究方向。基于Transformer的蛋白质大语言模型(LLMs)在蛋白序列设计和结构预测中展现潜力。大型抗体数据库如观测抗体空间(OAS)为抗体专用模型开发提供了可能。相比基于BERT的ProtBERT(420M参数),RoBERTa在性能更优的同时,参数量更小(125M),更适合实际部署。然而,基于RoBERTa的抗体专用模型尚未公开。本研究提出Ab-RoBERTa,一个基于RoBERTa的抗体专用大模型,已在Hugging Face公开(https://huggingface.co/mogam-ai/Ab-RoBERTa),可支持表位预测、人源性评估等应用。
原文摘要 · Abstract (English)
With the growing prominence of antibody-based therapeutics, antibody engineering has gained increasing attention as a critical area of research and development. Recent progress in transformer-based protein large language models (LLMs) has demonstrated promising applications in protein sequence design and structural prediction. Moreover, the availability of large-scale antibody datasets such as the Observed Antibody Space (OAS) database has opened new avenues for the development of LLMs specialized for processing antibody sequences. Among these, RoBERTa has demonstrated improved performance relative to BERT, while maintaining a smaller parameter count (125M) compared to the BERT-based protein model, ProtBERT (420M). This reduced model size enables more efficient deployment in antibody-related applications. However, despite the numerous advantages of the RoBERTa architecture, antibody-specific foundational models built upon it have remained inaccessible to the research community. In this study, we introduce Ab-RoBERTa, a RoBERTa-based antibody-specific LLM, which is publicly available at https://huggingface.co/mogam-ai/Ab-RoBERTa. This resource is intended to support a wide range of antibody-related research applications including paratope prediction or humanness assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。