arXiv:2606.28659q-bio.BMcs.LG2026-06

用Transformer主动学习,少数据高效筛选猪瘟疫苗候选抗原

Transformer-Based Active Learning for Data-Efficient Vaccine Epitope Selection in PRRS

  • 基于主动学习框架,用Transformer模型从少量数据中筛选高亲和力抗原
  • 仅30个样本时性能超传统方法用60个样本的效果,60个样本时准确率达86.8%
  • 适合低数据场景的疫苗设计,尤其对计算成本高的分子对接任务有实用价值

高保真分子对接模拟能有效预测抗原-受体结合亲和力,但计算成本高昂,限制了候选抗原的筛选数量。本文研究机器学习方法,利用主动学习策略在猪繁殖与呼吸综合征(PRRS)背景下,分类9聚体抗原与保守的猪白细胞抗原(SLA)受体之间的高亲和力结合。使用内部生成的80组抗原-SLA对接亲和力数据集,每组需超过48小时高性能计算(HPC)。在池基主动学习循环中,在严格低数据条件下训练多种模型(线性、MLP、CNN、小型Transformer)。通过大规模超参数优化,综合考虑模型架构、训练配置、获取策略和集成决策规则,确定最优配置。为减少数据子采样影响,每个配置均在多个随机且平衡的训练/验证子集上平均评估性能。实验表明,基于Transformer的序列模型始终表现最佳,主动增量学习显著优于随机采样基线。在中等数据量(N=30)下,优化模型性能超越用两倍数据训练的标准基线;在更高数据量(N=60)下,达到86.8%的峰值准确率,与基于构象噪声的两个独立估计上限85%一致。

原文摘要 · Abstract (English)

High-fidelity molecular docking simulations can produce biologically relevant estimates of epitope-receptor binding affinity but are computationally expensive and therefore limit the number of candidates that can be screened for vaccine design. In this work, we evaluate machine learning (ML) approaches where variants of active learning are used to classify instances of high binding affinity between 9-mer epitopes and a well-conserved swine leukocyte antigen (SLA) receptor in the context of Porcine Reproductive and Respiratory Syndrome (PRRS). We use an internally generated dataset of 80 epitope-SLA docking affinities, each requiring more than 48 hours of high-performance computing (HPC). Multiple model families (linear, MLP, CNN, and a small transformer) are trained under strict low-data conditions within a pool-based active learning loop. In each case, optimal model configurations are identified by conducting large-scale hyperparameter optimization over the combined space of model architecture, training configuration, acquisition policy, and ensemble decision rules. To mitigate the effects of data subsample selection, each candidate configuration is evaluated by averaging performance over many randomized and balanced training and validation data subsets. Across experiments, transformer-based sequence models consistently emerged as the best-performing architecture, with active incremental learning yielding significant improvement over a baseline random sample acquisition strategy. Under moderate training data availability (N=30), the optimized ML-model configuration outperforms a standard baseline trained on twice the amount of data. Under higher training data availability (N=60), the same configuration achieves a peak accuracy of 86.8%, consistent with an upper bound of 85% classification accuracy based on two independent estimates of conformational noise.

疫苗设计主动学习Transformer低数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。