首个纳米抗体综合评测基准,助力精准评估模型性能。
NbBench: Benchmarking Language Models for Comprehensive Nanobody Tasks
- 构建涵盖9个数据集的8类任务评测体系,覆盖结构、结合与开发性评估。
- 抗体语言模型在抗原相关任务表现优异,但对热稳定性等回归任务仍难突破。
- 首次统一标准,适合从事纳米抗体建模的研究者使用。
纳米抗体——源自骆驼科动物重链仅有的抗体片段——因其体积小、稳定性高、结合力强,在治疗和诊断中具有独特优势。尽管预训练蛋白质与抗体语言模型(PPLMs、PALMs)显著提升了生物分子理解能力,但针对纳米抗体的建模仍缺乏系统性研究与统一评测基准。为此,我们提出NbBench,首个面向纳米抗体表征学习的综合性基准。该基准涵盖九个精选数据集上的八类生物学意义任务,包括结构注释、结合亲和力预测及可开发性评估。我们系统评估了十一个代表性模型(含通用蛋白语言模型、抗体专用模型及纳米抗体专用模型),在冻结参数设置下进行测试。结果表明:抗体语言模型在抗原相关任务上表现突出,但在热稳定性、结合亲和力等回归任务上,所有模型均面临挑战,且无单一模型在所有任务中持续领先。通过统一数据集、任务定义与评估流程,NbBench为纳米抗体建模的可复现评估与持续进步提供了坚实基础。
原文摘要 · Abstract (English)
Nanobodies -- single-domain antibody fragments derived from camelid heavy-chain-only antibodies -- exhibit unique advantages such as compact size, high stability, and strong binding affinity, making them valuable tools in therapeutics and diagnostics. While recent advances in pretrained protein and antibody language models (PPLMs and PALMs) have greatly enhanced biomolecular understanding, nanobody-specific modeling remains underexplored and lacks a unified benchmark. To address this gap, we introduce NbBench, the first comprehensive benchmark suite for nanobody representation learning. Spanning eight biologically meaningful tasks across nine curated datasets, NbBench encompasses structure annotation, binding prediction, and developability assessment. We systematically evaluate eleven representative models -- including general-purpose protein LMs, antibody-specific LMs, and nanobody-specific LMs -- in a frozen setting. Our analysis reveals that antibody language models excel in antigen-related tasks, while performance on regression tasks such as thermostability and affinity remains challenging across all models. Notably, no single model consistently outperforms others across all tasks. By standardizing datasets, task definitions, and evaluation protocols, NbBench offers a reproducible foundation for assessing and advancing nanobody modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。