同时预测百日咳疫苗反应峰值与持久性,提升免疫评估精度。
Multitask Multimodal Fusion with Tabular Foundation Models for Peak and Durability Prediction of Pertussis Booster Response
- 多任务对比融合架构,联合建模峰值与持久性响应
- 在158人数据集上,峰值与持久性预测AUC分别达0.797和0.755
- 揭示了细胞因子与抗体特征对不同阶段预测的关键作用
百日咳加强针疫苗的免疫反应在个体间存在显著差异,表现为峰值强度与长期持久性两个阶段。这两个阶段由部分不同的生物学机制调控:峰值反映急性B细胞激活与抗体分泌,持久性则反映长期体液记忆建立。然而,现有计算模型通常仅关注单一阶段,忽略了完整的“激发-衰减”轨迹。联合预测两者具有挑战性,因两终点生物学解耦而非冗余;样本量小、多模态异构且存在结构化缺失,且两任务依赖不同测量时间窗。本文提出一种多任务对比多模态融合架构,结合冻结的TabPFN-v2模态编码器、双标签监督对比损失(若两受试者在任务1或任务2标签上一致,则视为正样本对)、基于实际缺失率校准的模态丢弃策略,以及缺失掩码注意力融合机制。应用于经筛选的CMI-PB百日咳加强针数据集(n = 158,四模态,44.9%样本至少缺一模态;峰值与持久性间斯皮尔曼相关系数r = -0.58,n = 96),模型在测试集上峰值预测AUROC为0.797(95% CI [0.621, 0.948]),持久性为0.755(95% CI [0.519, 0.945]),且在联合标签置换检验中均显著(N = 1000;p = 0.002,p = 0.045)。在原始特征与TabPFN嵌入上的逻辑回归、XGBoost、MLP基线比较中,本模型是唯一一个在两项任务上95%置信区间均高于随机水平的模型。各模态贡献分析结果与免疫学机制一致:峰值预测主要依赖细胞因子特征,持久性则依赖基线抗体特征。
原文摘要 · Abstract (English)
Pertussis booster vaccination produces immune responses that vary widely across individuals in both peak magnitude and long-term durability. These two phases are governed by partly distinct biological compartments:peak reflects acute B-cell activation and antibody secretion, while durability reflects the establishment of long-term humoral memory. Yet most computational models target only one, missing the full boost-and-wane trajectory. Jointly predicting both is non-trivial because the two endpoints are biologically dissociated rather than redundant; samples are small, modalities are heterogeneous with structured missingness, and the two tasks rely on different measurement windows. We propose a multi-task contrastive multimodal fusion architecture combining frozen TabPFN-v2 per-modality encoders, a dual-label supervised contrastive loss that treats two subjects as a positive pair if they agree on the Task 1 label or the Task 2 label, modality dropout calibrated to empirical missingness, and missingness-masked attention fusion. Applied to a curated subset of the CMI-PB pertussis booster dataset (n = 158 subjects, four modalities, 44.9% with at least one modality missing; Spearman r = -0.58 between peak and durability, n = 96), the model achieves test AUROC 0.797 (95% CI [0.621, 0.948]) for peak response and 0.755 (95% CI [0.519, 0.945]) for durability, with both significant under joint label permutation (N = 1000; p = 0.002 and p = 0.045). Across logistic regression, XGBoost, and MLP baselines on raw features and on TabPFN embeddings, the proposed model is the only one whose 95% CIs lie above chance on both tasks simultaneously. Per-modality contribution analyses recover task-specific modality contributions consistent with the underlying immunology: peak prediction is carried by cytokine signatures, while durability is carried by baseline antibody features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。