arXiv:2507.07839eess.IVcs.CV2025-07被引 16

融合影像、病理、基因等多模态数据,提升肾癌复发预测准确率。

MeD-3D: A Multimodal Deep Learning Framework for Precise Recurrence Prediction in Clear Cell Renal Cell Carcinoma (ccRCC)

  • 用多模态深度学习整合影像、病理、基因和临床数据
  • 在多个数据集上复发预测准确率达82.6%,优于单一模态模型
  • 支持部分数据缺失时推理,适合真实临床场景

透明细胞肾细胞癌(ccRCC)的复发预测因分子、病理和临床异质性而极具挑战。传统基于单一数据模态(如影像、病理或基因组)的预后模型难以全面捕捉疾病复杂性,导致预测效果不佳。本研究提出一种多模态深度学习框架MeD-3D,融合CT、MRI、组织病理全切片图像(WSI)、临床数据及基因组信息,以提升ccRCC复发预测精度并支持临床决策。该框架基于TCGA、TCIA和CPTAC等公开数据集构建,采用领域专用模型:基于ResNet50的CLAM处理病理WSI,预训练3D-ResNet18的MeD-3D处理影像数据,多层感知机(MLP)处理结构化临床与基因数据。各模态提取深层特征嵌入后,通过早期与晚期融合策略实现互补信息整合。框架还具备处理不完整数据的能力,在部分模态缺失时仍可进行有效推理,更贴合真实临床环境。

原文摘要 · Abstract (English)

Accurate prediction of recurrence in clear cell renal cell carcinoma (ccRCC) remains a major clinical challenge due to the disease complex molecular, pathological, and clinical heterogeneity. Traditional prognostic models, which rely on single data modalities such as radiology, histopathology, or genomics, often fail to capture the full spectrum of disease complexity, resulting in suboptimal predictive accuracy. This study aims to overcome these limitations by proposing a deep learning (DL) framework that integrates multimodal data, including CT, MRI, histopathology whole slide images (WSI), clinical data, and genomic profiles, to improve the prediction of ccRCC recurrence and enhance clinical decision-making. The proposed framework utilizes a comprehensive dataset curated from multiple publicly available sources, including TCGA, TCIA, and CPTAC. To process the diverse modalities, domain-specific models are employed: CLAM, a ResNet50-based model, is used for histopathology WSIs, while MeD-3D, a pre-trained 3D-ResNet18 model, processes CT and MRI images. For structured clinical and genomic data, a multi-layer perceptron (MLP) is used. These models are designed to extract deep feature embeddings from each modality, which are then fused through an early and late integration architecture. This fusion strategy enables the model to combine complementary information from multiple sources. Additionally, the framework is designed to handle incomplete data, a common challenge in clinical settings, by enabling inference even when certain modalities are missing.

多模态学习癌症预测深度学习影像病理融合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。