用预训练蛋白模型+半监督学习,解决流感疫苗更新慢的标签数据瓶颈。
Mitigating the Antigenic Data Bottleneck: Semi-supervised Learning with Protein Language Models for Influenza A Surveillance
- 结合蛋白语言模型与半监督学习,提升低标注数据下的抗原性预测精度。
- 仅用25%标注数据时,ESM-2模型仍保持0.82以上F1分数。
- 特别适合流感病毒监测中标签稀缺但序列海量的场景。
流感A病毒(IAVs)抗原变异速度快,需频繁更新疫苗,但用于量化抗原性的血凝抑制(HI)实验耗时且难扩展,导致基因组数据远超可获得的表型标签,限制了传统监督模型效果。我们提出将预训练蛋白语言模型(PLMs)与半监督学习(SSL)结合,在标签稀缺情况下仍保持高预测准确率。评估了自训练和标签传播两种SSL策略,使用四种PLM嵌入(ESM-2、ProtVec、ProtT5、ProtBert)处理血凝素(HA)序列,在四个亚型(H1N1、H3N2、H5N1、H9N2)上采用嵌套交叉验证模拟低标签场景(25%、50%、75%、100%)。结果表明,SSL在标签稀缺时持续提升性能;使用ProtVec的自训练取得最大相对增益,说明其可补偿低分辨率表示;ESM-2表现最稳健,25%标签下F1分数仍超0.82,表明其嵌入捕获了关键抗原决定簇。尽管H1N1和H9N2预测准确率高,但高变异性亚型H3N2仍具挑战,但SSL有效缓解了性能下降。研究证明,融合PLMs与SSL可突破抗原性标注瓶颈,更高效利用未标注监测序列,支持快速变异株优先排序与及时疫苗株选择。
原文摘要 · Abstract (English)
Influenza A viruses (IAVs) evolve antigenically at a pace that requires frequent vaccine updates, yet the haemagglutination inhibition (HI) assays used to quantify antigenicity are labor-intensive and unscalable. As a result, genomic data vastly outpace available phenotypic labels, limiting the effectiveness of traditional supervised models. We hypothesize that combining pre-trained Protein Language Models (PLMs) with Semi-Supervised Learning (SSL) can retain high predictive accuracy even when labeled data are scarce. We evaluated two SSL strategies, Self-training and Label Spreading, against fully supervised baselines using four PLM-derived embeddings (ESM-2, ProtVec, ProtT5, ProtBert) applied to haemagglutinin (HA) sequences. A nested cross-validation framework simulated low-label regimes (25%, 50%, 75%, and 100% label availability) across four IAV subtypes (H1N1, H3N2, H5N1, H9N2). SSL consistently improved performance under label scarcity. Self-training with ProtVec produced the largest relative gains, showing that SSL can compensate for lower-resolution representations. ESM-2 remained highly robust, achieving F1 scores above 0.82 with only 25% labeled data, indicating that its embeddings capture key antigenic determinants. While H1N1 and H9N2 were predicted with high accuracy, the hypervariable H3N2 subtype remained challenging, although SSL mitigated the performance decline. These findings demonstrate that integrating PLMs with SSL can address the antigenicity labeling bottleneck and enable more effective use of unlabeled surveillance sequences, supporting rapid variant prioritization and timely vaccine strain selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。