arXiv:2601.20347cs.CV2026-01被引 1

融合病理图像与临床数据,提升癌症分类与生存预测精度

MMSF: Multitask and Multimodal Supervised Framework for WSI Classification and Survival Analysis

  • 构建多任务多模态框架,分步融合组织拓扑与患者临床信息
  • 在多个数据集上分类准确率提升2.1%~6.6%,生存预测C指数提升7.1%~9.8%
  • 适合做癌症精准诊疗、多模态医学影像分析的研究者参考

多模态证据对计算病理学至关重要:超像素全切片图像捕捉肿瘤形态,而患者级临床描述提供预后互补信息。由于特征空间统计差异大、尺度不一,整合此类异构信号仍具挑战。我们提出MMSF,一种基于线性复杂度的MIL骨干网络的多任务多模态监督框架,显式分解并融合跨模态信息。MMSF包含图特征提取模块(嵌入切片级别组织拓扑)、临床数据嵌入模块(标准化患者属性)、特征融合模块(对齐共模态与特异性表示),以及基于Mamba的MIL编码器和多任务预测头。在CAMELYON16和TCGA-NSCLC上的实验表明,相比竞争基线,分类准确率提升2.1%~6.6%,AUC提升2.2%~6.9%;在五个TCGA生存队列评估中,相比单模态方法C-index提升7.1%~9.8%,相比多模态方法提升5.6%~7.1%。

原文摘要 · Abstract (English)

Multimodal evidence is critical in computational pathology: gigapixel whole slide images capture tumor morphology, while patient-level clinical descriptors preserve complementary context for prognosis. Integrating such heterogeneous signals remains challenging because feature spaces exhibit distinct statistics and scales. We introduce MMSF, a multitask and multimodal supervised framework built on a linear-complexity MIL backbone that explicitly decomposes and fuses cross-modal information. MMSF comprises a graph feature extraction module embedding tissue topology at the patch level, a clinical data embedding module standardizing patient attributes, a feature fusion module aligning modality-shared and modality-specific representations, and a Mamba-based MIL encoder with multitask prediction heads. Experiments on CAMELYON16 and TCGA-NSCLC demonstrate 2.1--6.6\% accuracy and 2.2--6.9\% AUC improvements over competitive baselines, while evaluations on five TCGA survival cohorts yield 7.1--9.8\% C-index improvements compared with unimodal methods and 5.6--7.1\% over multimodal alternatives.

多模态学习病理图像分析生存分析Mamba

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。