PRIME通过原型记忆实现缺失模态下的癌症预后多模态预训练。
PRIME: Prototype-Driven Multimodal Pretraining for Cancer Prognosis with Missing Modalities
- 用统一嵌入空间和原型记忆银行,实现跨模态语义补全。
- 在32种癌症上预训练,三任务平均C-index达0.653,优于现有方法。
- 适合临床数据缺失严重的场景,支持高效迁移与小样本适应。
多模态自监督预训练可通过整合组织病理全片图像、基因表达和病理报告来提升癌症预后预测能力,但现有方法通常要求输入完全配对且完整。实际临床队列常存在模态缺失,限制了监督融合与可扩展的多模态预训练。我们提出PRIME,一种面向缺失数据的多模态自监督预训练框架,能从部分观测队列中学习鲁棒且可迁移的表示。PRIME将异构模态嵌入映射到统一标记空间,并引入共享原型内存库,通过患者级共识检索实现潜在空间语义补全,生成结构对齐的标记而无需重建原始信号。两种互补的预训练目标——跨模态对齐与结构化缺失增强下的融合后一致性——联合学习在任意模态子集下仍具预测力的表示。我们在癌症基因组图谱(TCGA)上进行了无标签预训练,覆盖32种癌症,下游在五个队列上进行五折评估,涵盖总生存预测、三年死亡率分类和三年复发率分类。PRIME在三项任务中均取得最佳宏观平均表现,分别达到0.653 C-index、0.689 AUROC和0.637 AUROC,同时在测试时缺失情形下表现出更强鲁棒性,并支持参数高效与标签高效适配。结果表明,面向缺失的多模态预训练是碎片化临床数据环境中预后建模的实用策略。
原文摘要 · Abstract (English)
Multimodal self-supervised pretraining offers a promising route to cancer prognosis by integrating histopathology whole-slide images, gene expression, and pathology reports, yet most existing approaches require fully paired and complete inputs. In practice, clinical cohorts are fragmented and often miss one or more modalities, limiting both supervised fusion and scalable multimodal pretraining. We propose PRIME, a missing-aware multimodal self-supervised pretraining framework that learns robust and transferable representations from partially observed cohorts. PRIME maps heterogeneous modality embeddings into a unified token space and introduces a shared prototype memory bank for latent-space semantic imputation via patient-level consensus retrieval, producing structurally aligned tokens without reconstructing raw signals. Two complementary pretraining objectives: inter-modality alignment and post-fusion consistency under structured missingness augmentation, jointly learn representations that remain predictive under arbitrary modality subsets. We evaluate PRIME on The Cancer Genome Atlas with label-free pretraining on 32 cancer types and downstream 5-fold evaluation on five cohorts across overall survival prediction, 3-year mortality classification, and 3-year recurrence classification. PRIME achieves the best macro-average performance among all compared methods, reaching 0.653 C-index, 0.689 AUROC, and 0.637 AUROC on the three tasks, respectively, while improving robustness under test-time missingness and supporting parameter-efficient and label-efficient adaptation. These results support missing-aware multimodal pretraining as a practical strategy for prognosis modeling in fragmented clinical data settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。