融合病理图像与电子病历,提升乳腺癌早期诊断准确率
Multimodal Fusion of Histopathology Images and Electronic Health Records for Early Breast Cancer Diagnosis

- 将病理切片特征与临床数据在中间层融合,增强判别力
- 多模态模型在三分类任务上达到AUC 0.997,对关键指标(核分裂象)提升显著
- 适用于需要高精度与可解释性的临床辅助诊断场景
乳腺癌是全球癌症致死的主要原因,及时准确的诊断对改善生存率至关重要。尽管卷积神经网络(CNN)在病理图像分类中表现优异,机器学习模型在结构化电子健康记录(EHR)风险分层中也展现出价值,但现有研究大多孤立处理这些模态。本文提出一个系统性多模态框架,整合来自BreCaHAD数据集的病理切片级特征与MIMIC-IV中的结构化临床数据。我们训练并评估了单模态图像模型(简单CNN基线和使用迁移学习的ResNet-18)、单模态表格模型(XGBoost与多层感知机),以及一个在中间层拼接双模态潜在表示的融合模型。ResNet-18在三类切片级分类上达到近乎完美的准确率(1.000)与AUC(1.000),XGBoost在EHR预测任务中实现98%准确率。中间融合模型的宏平均AUC达0.997,优于所有单模态基线,在诊断关键但类别不平衡的核分裂象类别上提升最显著(AUC 0.994)。Grad-CAM与SHAP可解释性分析表明,模型决策符合既定病理与临床标准。结果表明,多模态融合在预测性能与临床透明性方面均有显著提升。
原文摘要 · Abstract (English)
Breast cancer is a leading cause of cancer-related mortality worldwide, and timely accurate diagnosis is critical to improving survival outcomes. While convolutional neural networks (CNNs) have demonstrated strong performance on histopathology image classification, and machine learning models on structured electronic health records (EHR) have shown utility for clinical risk stratification, most existing work treats these modalities in isolation. This paper presents a systematic multimodal framework that integrates patch-level histopathology features from the BreCaHAD dataset with structured clinical data from MIMIC-IV. We train and evaluate unimodal image models (a simple CNN baseline and ResNet-18 with transfer learning), unimodal tabular models (XGBoost and a multilayer perceptron), and an intermediate-fusion model that concatenates latent representations from both modalities. ResNet-18 achieves near-perfect accuracy (1.000) and AUC (1.000) on three-class patch-level classification, while XGBoost achieves 98% accuracy on the EHR prediction task. The intermediate fusion model yields a macro-average AUC of 0.997, outperforming all unimodal baselines and delivering the largest improvements on the diagnostically critical but class-imbalanced mitosis category (AUC 0.994). Grad-CAM and SHAP interpretability analyses validate that model decisions align with established pathological and clinical criteria. Our results demonstrate that multimodal integration delivers meaningful improvements in both predictive performance and clinical transparency.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。