用潜在空间预测增强蛋白语言模型,效果优于纯掩码建模。
ProteinJEPA: Latent prediction complements protein language models
- 仅在掩码位置预测潜在表示,结合掩码建模目标
- 16项下游任务中10~11胜3负,提升稳定性与远源同源识别能力
- 适合追求预训练效率与泛化性能的生物序列研究者
蛋白语言模型主要通过掩码语言建模(MLM)训练,预测被遮蔽位置的氨基酸。我们探究在相同计算时间内,潜在空间预测是否能补充这种逐标记目标。在35–150M参数的预训练与随机初始化蛋白序列编码器上,最佳的Protein-JEPA设计并非全位置潜在预测,而是仅在掩码位置预测潜在目标,并保留MLM交叉熵损失。该方法称为掩码位置MLM+JEPA。在16项下游任务(15个冻结线性探针加SCOPe-40零样本折叠检索)中,该方法在匹配计算时间预算下,对ESM2-35M实现10胜3负3平,对ESM2-150M实现11胜2负3平;从头预训练结果则参差不齐(6胜8负2平)。多个模型在11个任务中表现更优,包括稳定性、β-内酰胺酶适应度、变异效应、内在无序性、远源同源、酶分类和SCOPe-40折叠检索。失败较多的任务为荧光(TAPE)和肽-HLA结合。全位置MLM+JEPA整体持平但无法复现掩码位置优势,而纯JEPA几乎全部崩溃。结论:当与MLM结合时,JEPA具有竞争力,甚至可在匹配计算预算下超越纯MLM。
原文摘要 · Abstract (English)
Protein language models are trained primarily with masked language modeling (MLM), which predicts amino-acid identities at masked positions. We ask whether latent-space prediction can complement these token-level objectives under matched wall-clock budget. Across pretrained and random-init protein sequence encoders at 35--150M parameters, we find that the best protein-JEPA design is not all-position latent prediction but a variant: predicting latent targets only at masked positions, and retaining the MLM cross-entropy. We call this recipe masked-position MLM+JEPA. On a 16-task downstream suite (15 frozen linear probes plus SCOPe-40 zero-shot fold retrieval), under matched wall-clock budgets, this recipe wins more tasks than it loses against MLM-only continuation: 10 wins / 3 losses / 3 ties (hereafter W/L/T) on pretrained ESM2-35M, 11/2/3 on ESM2-150M while results in pretraining from scratch are mixed (6/8/2). Gains are seen for multiple models on 11 of 16 tasks, including stability, \b{eta}β\b{eta}-lactamase fitness, variant effect, intrinsic disorder, remote homology, enzyme classification, and SCOPe-40 fold retrieval. Tasks with more losses than wins are Fluorescence (TAPE) and Peptide-HLA Binding. All-position MLM+JEPA matches MLM-only overall but does not reproduce the masked-position gains. JEPA-only (no MLM) collapses in nearly every experiment. We conclude that JEPA, when combined with MLM, is competitive and can outperform pure MLM in pretraining and continued training, even under matched wall-clock budgets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。