arXiv:2506.09052cs.LGcs.AI2025-06被引 3

用Llama3架构预测抗体结合亲和力,精度超现有方法。

Llama-Affinity: A Predictive Antibody Antigen Binding Model Integrating Antibody Sequences with Llama3 Backbone Architecture

  • 基于Llama3架构融合抗体序列数据进行亲和力预测
  • 准确率0.964、AUC-ROC达0.9936,优于现有模型
  • 训练效率提升五倍,仅需0.46小时/次

抗体介导的免疫应答是机体抵御病原体、病毒等外来入侵的关键。抗体特异性结合并中和抗原的能力对维持免疫力至关重要。近年来,生物工程技术显著加速了治疗性抗体的开发,这类药物在癌症、SARS-CoV-2、自身免疫疾病及传染病治疗中表现出卓越疗效。传统实验测定亲和力耗时且成本高。随着人工智能发展,基于机器学习的计算方法革新了虚拟药物设计,特别是大语言模型(LLMs)在抗体表征中的应用,为抗体亲和力预测开辟新路径。本文提出一种新型抗体-抗原结合亲和力预测模型(LlamaAffinity),采用开源Llama3架构,并利用来自观测抗体空间(OAS)数据库的抗体序列数据。该方法在多个评估指标上显著优于现有最先进模型(AntiFormer、AntiBERTa、AntiBERTy):准确率0.9640,F1分数0.9643,精确率0.9702,召回率0.9586,AUC-ROC达0.9936。此外,该策略具备更高计算效率,平均累计训练时间仅为0.46小时,较以往研究降低五倍。

原文摘要 · Abstract (English)

Antibody-facilitated immune responses are central to the body's defense against pathogens, viruses, and other foreign invaders. The ability of antibodies to specifically bind and neutralize antigens is vital for maintaining immunity. Over the past few decades, bioengineering advancements have significantly accelerated therapeutic antibody development. These antibody-derived drugs have shown remarkable efficacy, particularly in treating cancer, SARS-CoV-2, autoimmune disorders, and infectious diseases. Traditionally, experimental methods for affinity measurement have been time-consuming and expensive. With the advent of artificial intelligence, in silico medicine has been revolutionized; recent developments in machine learning, particularly the use of large language models (LLMs) for representing antibodies, have opened up new avenues for AI-based design and improved affinity prediction. Herein, we present an advanced antibody-antigen binding affinity prediction model (LlamaAffinity), leveraging an open-source Llama 3 backbone and antibody sequence data sourced from the Observed Antibody Space (OAS) database. The proposed approach shows significant improvement over existing state-of-the-art (SOTA) methods (AntiFormer, AntiBERTa, AntiBERTy) across multiple evaluation metrics. Specifically, the model achieved an accuracy of 0.9640, an F1-score of 0.9643, a precision of 0.9702, a recall of 0.9586, and an AUC-ROC of 0.9936. Moreover, this strategy unveiled higher computational efficiency, with a five-fold average cumulative training time of only 0.46 hours, significantly lower than in previous studies.

抗体预测大模型亲和力AI制药

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。