arXiv:2605.25050stat.APcs.LG2026-05

解决临床肿瘤数据中多模态缺失问题,提升免疫治疗耐药预测精度

Multimodality Stacking with Blockwise missing values and application to the PIONeeR biomarkers study for prediction of resistance to immunotherapy

论文配图:Multimodality Stacking with Blockwise missing values and application to the PIONeeR biomarkers study for prediction of resistance to immunotherapy
图 1 · 摘自论文原文
  • 分步建模各数据模态,用交叉验证堆叠元学习器融合预测结果
  • 在443名患者上提升生存预测效果,线性模型提升15.9%(p<0.001)
  • 可处理完整缺失数据,适合生物标志物系统评估与医学研究应用

临床肿瘤学中整合多模态数据常受限于高维特征与块状缺失问题,即特定患者群体缺乏完整数据源。传统生存分析模型难以应对此类缺失,易导致偏差或排除患者。本文提出多模态堆叠框架MSB,通过独立建模各模态特征,并利用交叉验证的堆叠元学习器聚合预测结果。在包含443名患者、378个生物标志物及八个异构数据源的PIONeeR研究中,用于预测接受免疫治疗的晚期非小细胞肺癌患者的无进展生存期。结果表明,相较于基线算法,MSB显著提升预测性能(C-index):线性模型提高15.9%(p<0.001),随机生存森林提升5.4%(p=0.002),梯度提升方法提升2.1%(p=0.030)。此外,MSB降低泛化差距(训练-测试差异:0.055 vs 0.380,5折交叉验证重复3次)。置换重要性分析显示,常规实验室指标、临床特征和PD-L1表达为主要预测因子。缺失块指示符重要性极低,说明模型依赖于生物标志物值而非数据可用性模式。该框架为存在块状缺失的多模态生存预测提供了统计验证方案,支持无需完整数据即可系统评估生物标志物,具有实际应用价值,需外部验证。代码已开源,地址:https://github.com/MohamedBoussena/MSB,采用Inria许可证。

原文摘要 · Abstract (English)

Integrating multimodal datasets in clinical oncology is frequently hindered by high dimensionality and blockwise missingness, where entire data sources are unavailable for specific patient subsets. Standard survival models often struggle with these gaps, leading to biased results or patient exclusion. We introduce Multimodality Stacking with Blockwise missing values (MSB), a late-fusion framework for survival analysis that independently models modality-specific features before aggregating predictions via a cross-validated stacking meta-learner. MSB was validated on the PIONeeR study (n=443 patients, 378 biomarkers across eight heterogeneous sources) to predict progression-free survival in advanced non-small cell lung cancer patients receiving immunotherapy. MSB yielded higher predictive performance (C-index) than baseline algorithms. Improvements varied by baseline strength: linear models showed a 15.9% increase (p<0.001 for the Wilcoxon signed-rank test), random survival forests gained 5.4% (p=0.002), and gradient boosting methods improved by 2.1% (p=0.030). Beyond discrimination, MSB reduced the generalization gap (train-test difference in 5 folds cross-validation repeated 3 times: 0.055 vs 0.380 for linear models). Permutation importance analysis identified routine laboratory markers, clinical features, and PD-L1 expression as primary predictive drivers. Missing block indicators showed negligible importance, suggesting the model learned from biomarker values rather than data availability patterns. MSB provides a statistically validated framework for multimodal survival prediction with blockwise missingness. By enabling systematic biomarker evaluation without requiring complete data, MSB offers a practical tool for predictive modeling in biomedical research, pending external validation. Implementation is available at https://github.com/MohamedBoussena/MSB under Inria license.

生存分析多模态融合缺失数据生物标志物

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。