arXiv:2605.12852cs.LGq-bio.QM2026-05

同时预测百日咳疫苗反应峰值与持久性,提升免疫评估精度。

Multitask Multimodal Fusion with Tabular Foundation Models for Peak and Durability Prediction of Pertussis Booster Response

  • 多任务对比融合架构,联合建模峰值与持久性响应
  • 在158人数据集上,峰值与持久性预测AUC分别达0.797和0.755
  • 揭示了细胞因子与抗体特征对不同阶段预测的关键作用

百日咳加强针疫苗的免疫反应在个体间存在显著差异,表现为峰值强度与长期持久性两个阶段。这两个阶段由部分不同的生物学机制调控:峰值反映急性B细胞激活与抗体分泌,持久性则反映长期体液记忆建立。然而,现有计算模型通常仅关注单一阶段,忽略了完整的“激发-衰减”轨迹。联合预测两者具有挑战性,因两终点生物学解耦而非冗余;样本量小、多模态异构且存在结构化缺失,且两任务依赖不同测量时间窗。本文提出一种多任务对比多模态融合架构,结合冻结的TabPFN-v2模态编码器、双标签监督对比损失(若两受试者在任务1或任务2标签上一致,则视为正样本对)、基于实际缺失率校准的模态丢弃策略,以及缺失掩码注意力融合机制。应用于经筛选的CMI-PB百日咳加强针数据集(n = 158,四模态,44.9%样本至少缺一模态;峰值与持久性间斯皮尔曼相关系数r = -0.58,n = 96),模型在测试集上峰值预测AUROC为0.797(95% CI [0.621, 0.948]),持久性为0.755(95% CI [0.519, 0.945]),且在联合标签置换检验中均显著(N = 1000;p = 0.002,p = 0.045)。在原始特征与TabPFN嵌入上的逻辑回归、XGBoost、MLP基线比较中,本模型是唯一一个在两项任务上95%置信区间均高于随机水平的模型。各模态贡献分析结果与免疫学机制一致:峰值预测主要依赖细胞因子特征,持久性则依赖基线抗体特征。

原文摘要 · Abstract (English)

Pertussis booster vaccination produces immune responses that vary widely across individuals in both peak magnitude and long-term durability. These two phases are governed by partly distinct biological compartments:peak reflects acute B-cell activation and antibody secretion, while durability reflects the establishment of long-term humoral memory. Yet most computational models target only one, missing the full boost-and-wane trajectory. Jointly predicting both is non-trivial because the two endpoints are biologically dissociated rather than redundant; samples are small, modalities are heterogeneous with structured missingness, and the two tasks rely on different measurement windows. We propose a multi-task contrastive multimodal fusion architecture combining frozen TabPFN-v2 per-modality encoders, a dual-label supervised contrastive loss that treats two subjects as a positive pair if they agree on the Task 1 label or the Task 2 label, modality dropout calibrated to empirical missingness, and missingness-masked attention fusion. Applied to a curated subset of the CMI-PB pertussis booster dataset (n = 158 subjects, four modalities, 44.9% with at least one modality missing; Spearman r = -0.58 between peak and durability, n = 96), the model achieves test AUROC 0.797 (95% CI [0.621, 0.948]) for peak response and 0.755 (95% CI [0.519, 0.945]) for durability, with both significant under joint label permutation (N = 1000; p = 0.002 and p = 0.045). Across logistic regression, XGBoost, and MLP baselines on raw features and on TabPFN embeddings, the proposed model is the only one whose 95% CIs lie above chance on both tasks simultaneously. Per-modality contribution analyses recover task-specific modality contributions consistent with the underlying immunology: peak prediction is carried by cytokine signatures, while durability is carried by baseline antibody features.

多任务学习免疫预测多模态融合生物医学建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。