arXiv:2603.00253cs.LG2026-03

构建蛋白语言模型持续预训练基准,评估方法在十年数据上的表现。

CoPeP: Benchmarking Continual Pretraining for Protein Language Models

  • 基于十年UniProt数据序列构建持续学习评估基准
  • 利用时间元信息使困惑度降低最多7%
  • 适合研究持续学习与生物序列建模的学者参考

蛋白质语言模型(pLMs)近年来受到广泛关注,因其能从进化统计中揭示序列、结构与功能间的关系,从而加速药物发现。这些模型依赖于生物学界持续更新的大规模蛋白数据库,其动态特性促使采用持续学习,不仅应对数据增长,还可利用过程中的时间元信息。为此,我们提出持续预训练蛋白语言模型(CoPeP)基准,首次系统评估持续学习方法在pLMs上的表现。具体而言,我们从UniProt知识库中提取跨越十年的蛋白数据集,并定义31项蛋白理解任务的评估指标。我们测试了包括回放、遗忘和可塑性方法在内的多种持续学习策略,部分方法此前未在如此大规模的模型与数据上应用。结果表明,引入时间元信息可使困惑度降低达7%,即便相比联合训练所有任务数据仍具优势。即使在大规模场景下,部分持续学习方法仍优于简单持续预训练。CoPeP基准为在真实世界应用中大规模研究这些方法提供了重要机会。

原文摘要 · Abstract (English)

Protein language models (pLMs) have recently gained significant attention for their ability to uncover relationships between sequence, structure, and function from evolutionary statistics, thereby accelerating therapeutic drug discovery. These models learn from large protein databases that are continuously updated by the biology community and whose dynamic nature motivates the application of continual learning, not only to keep up with the ever-growing data, but also as an opportunity to take advantage of the temporal meta-information that is created during this process. As a result, we introduce the Continual Pretraining of Protein Language Models (CoPeP) benchmark, a novel benchmark for evaluating continual learning approaches on pLMs. Specifically, we curate a sequence of protein datasets derived from the UniProt Knowledgebase spanning a decade and define metrics to assess pLM performance across 31 protein understanding tasks. We evaluate several methods from the continual learning literature, including replay, unlearning, and plasticity-based methods, some of which have never been applied to models and data of this scale. Our findings reveal that incorporating temporal meta-information improves perplexity by up to 7% even when compared to training on data from all tasks jointly. Moreover, even at scale, several continual learning methods outperform naive continual pretraining. The CoPeP benchmark offers an exciting opportunity to study these methods at scale in an impactful real-world application.

蛋白语言模型持续学习基准评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。