解耦肿瘤与微环境,实现病理与基因数据高效融合诊断癌症
Disentangled Multi-modal Learning of Histology and Transcriptomics for Cancer Characterization
- 将病理图像与基因数据解耦为肿瘤和微环境两部分,分别建模
- 在无配对数据情况下仍可精准预测癌症预后,准确率超现有方法
- 适合病理与基因联合分析的研究者,尤其关注癌症分型与预后者
组织病理学仍是癌症诊断与预后的金标准。随着转录组测序的发展,结合病理图像与转录组的多模态学习能提供更全面的信息。然而,现有方法受限于模态间固有异质性、多尺度信息整合不足及依赖成对数据,制约了临床应用。为此,我们提出一种解耦式多模态框架,包含四项贡献:1)通过解耦多模态融合模块,将全切片图像(WSI)与转录组分解为肿瘤与微环境子空间,并引入置信度引导的梯度协调策略平衡子空间优化;2)提出跨放大倍数基因表达一致性策略,增强不同分辨率下的转录组信号对齐;3)设计子空间知识蒸馏策略,使仅依赖病理图像的模型也能实现无转录组推理;4)提出有信息量的标记聚合模块,抑制图像冗余同时保留子空间语义。在癌症诊断、预后与生存预测任务中,本方法在多种设置下均优于现有先进方法。代码已开源:https://github.com/helenypzhang/Disentangled-Multimodal-Learning。
原文摘要 · Abstract (English)
Histopathology remains the gold standard for cancer diagnosis and prognosis. With the advent of transcriptome profiling, multi-modal learning combining transcriptomics with histology offers more comprehensive information. However, existing multi-modal approaches are challenged by intrinsic multi-modal heterogeneity, insufficient multi-scale integration, and reliance on paired data, restricting clinical applicability. To address these challenges, we propose a disentangled multi-modal framework with four contributions: 1) To mitigate multi-modal heterogeneity, we decompose WSIs and transcriptomes into tumor and microenvironment subspaces using a disentangled multi-modal fusion module, and introduce a confidence-guided gradient coordination strategy to balance subspace optimization. 2) To enhance multi-scale integration, we propose an inter-magnification gene-expression consistency strategy that aligns transcriptomic signals across WSI magnifications. 3) To reduce dependency on paired data, we propose a subspace knowledge distillation strategy enabling transcriptome-agnostic inference through a WSI-only student model. 4) To improve inference efficiency, we propose an informative token aggregation module that suppresses WSI redundancy while preserving subspace semantics. Extensive experiments on cancer diagnosis, prognosis, and survival prediction demonstrate our superiority over state-of-the-art methods across multiple settings. Code is available at https://github.com/helenypzhang/Disentangled-Multimodal-Learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。