arXiv:2608.06022cs.CLq-bio.GN2026-08

测试大模型能否从序列中理解抗原表位,助力抗体药物研发。

EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?

论文配图:EpiBench: Can LLMs Understand Epitopes for Antibody Drug Discovery?
图 1 · 摘自论文原文
  • 构建基于序列的闭卷评测基准,覆盖五类表位相关任务。
  • 1609个样本涵盖结构接触与功能实验数据,支持精准评估。
  • 揭示当前大模型在抗体序列定位与生物推理上的不足。

表位决定抗体结合抗原的位置,影响治疗效果如功能阻断和逃逸耐药性,是抗体药物研发的核心。尽管大语言模型(LLMs)展现出强大的生物医学推理能力,但其是否能直接从抗原与抗体序列中推断表位信息仍不明确。现有表位资源多聚焦单一预测任务或依赖特殊结构设定,而通用蛋白质基准又未评估抗体开发流程中的表位中心决策。为此,我们提出EpiBench,一个闭卷、基于序列且可自动评分的基准,用于评估LLMs在表位推理方面的能力。EpiBench包含1,609个基于结构抗体-抗原接触、功能B细胞检测及深度突变扫描逃逸测量的精心标注样本,覆盖五项关联任务:可靶向区域发现、抗体条件下的表位识别、表位分组、功能表位评估和抗体逃逸评估,并通过受控采样减少捷径学习带来的评估偏差。我们评估了九个通用型LLMs,并通过任务特异性基线、抗原长度分层、显式推理对比及失败模式分析其表现。结果表明,当前LLMs虽能捕捉部分表位相关信号,但在抗体特异性序列锚定、长上下文残基定位和生物学合理推理方面仍存在局限。因此,EpiBench为衡量和提升序列感知型生物医学LLMs提供了诊断性测试平台,推动可靠的大模型辅助抗体发现。

原文摘要 · Abstract (English)

Epitopes determine where antibodies bind antigens and shape downstream therapeutic properties such as functional blockade and escape resistance, making epitope understanding central to antibody drug discovery. Although large language models (LLMs) have shown strong biomedical reasoning ability, it remains unclear whether they can infer epitope information directly from antigen and antibody sequences. Existing epitope resources typically focus on isolated prediction tasks or rely on specialized structural settings, while general protein benchmarks do not evaluate epitope-centered decisions across the antibody development workflow. To address this gap, we introduce EpiBench, a closed-book, sequence-based, and automatically scorable benchmark for evaluating epitope reasoning in LLMs. EpiBench contains 1,609 curated samples grounded in structural antibody--antigen contacts, curated functional B-cell assays, and deep mutational scanning escape measurements. It covers five connected tasks: targetable region discovery, antibody-conditioned epitope identification, epitope binning, functional epitope assessment, and antibody escape assessment, with controlled sampling to reduce shortcut-based evaluation artifacts. We evaluate nine general-purpose LLMs and analyze their behavior through task-specific baselines, antigen length stratification, explicit-reasoning comparison, and failure-mode inspection. The results show that current LLMs capture partial epitope-related signals but remain limited in antibody-specific sequence grounding, long-context residue localization, and biologically grounded reasoning. Therefore, EpiBench provides a diagnostic testbed for measuring and improving sequence-aware biomedical LLMs toward reliable LLM-assisted antibody discovery.

表位识别抗体药物大模型评测生物医学

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。