构建病毒蛋白突变评估基准,验证大模型预测变异影响能力
ViroGym: Realistic Large-Scale Benchmarks for Evaluating Viral Proteins
- 设计三类任务:55万+突变序列的深度突变扫描、流感中和实验、新冠疫情预测
- ProGen2家族在所有任务中表现最优,且体外测试结果可预测真实突变演化
- 不同实验数据集虽突变重叠少,但能互补反映进化约束,适合疫苗与药物研发
蛋白质语言模型(pLMs)在零样本预测错义突变效应方面展现出强大潜力,但针对病毒蛋白的系统性评估仍显不足,这在需要提前预警新突变的背景下尤为关键。本文提出ViroGym,一个涵盖三类任务的综合性基准:79项深度突变扫描(DMS)实验,覆盖真核病毒,包含552,065个突变序列及7种表型读数;21项流感中和任务;以及一项针对SARS-CoV-2的真实世界疫情预测任务。我们对多个经典pLMs在适应度景观、抗原多样性与大流行预测能力上进行评估,发现ProGen2家族在所有任务中表现最佳。关键的是,DMS与中和实验性能可稳定预测模型在真实突变演化中的泛化能力,尽管其突变集合几乎无重叠,表明互补的体外基准能有效捕捉真实突变预测所需的进化约束。
原文摘要 · Abstract (English)
Protein language models (pLMs) have shown strong potential for zero-shot prediction of missense variant effects, yet systematic benchmarking on viral proteins remains limited, a critical gap given the need for proactive tools that can anticipate emerging mutations ahead of experimental validation. Here we introduce ViroGym, a comprehensive benchmark evaluating pLMs across three tasks: 79 deep mutational scanning (DMS) assays covering eukaryotic viruses with 552,065 mutated sequences across 7 phenotypic readouts, 21 influenza neutralisation tasks, and a real-world pandemic prediction task for SARS-CoV-2. We benchmark well-established pLMs on fitness landscapes, antigenic diversity, and pandemic forecasting, and find that the ProGen2 family consistently achieves the strongest performance across all three tasks. Crucially, DMS and neutralisation performance reliably identifies models that generalise to real-world emergence, even though the mutation sets they surface barely overlap, revealing that complementary in vitro benchmarks capture the evolutionary constraints needed for real-world mutation forecasting.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。