用多模态机器学习提前一个月预测四种癌症转移风险,表现优于传统方法。
Multimodal Machine Learning for Early Prediction of Metastasis in a Swedish Multi-Cancer Cohort
- 融合电子病历中的文本、检验、用药等多源数据,采用中间融合策略提升预测能力。
- 在乳腺、前列腺癌中F1得分达0.845,肺癌文本模型最佳达0.829,整体表现优异。
- 适合临床早期预警系统开发,尤其关注多模态数据整合与可解释性分析的团队。
多模态机器学习通过整合电子健康记录(EHR)中的结构化与非结构化数据,提供患者状态的全景视图。本文提出一种框架,利用患者六个月内临床历史数据,提前一个月预测转移风险。分析了瑞典卡罗林斯卡大学医院的四个癌症队列:乳腺癌(n=743)、结肠癌(n=387)、肺癌(n=870)和前列腺癌(n=1890),数据涵盖人口学、共病、实验室结果、药物及临床文本。对比了传统与深度学习分类器在单模态与多模态组合下的表现,采用多种融合策略,并遵循TRIPOD 2a设计,以80-20划分训练验证集确保评估严谨性。性能指标包括AUROC、AUPRC、F1分数、敏感性和特异性。进一步使用改进的SHAP方法分析模型决策逻辑。中间融合在乳腺癌(F1=0.845)、结肠癌(F1=0.786)和前列腺癌(F1=0.845)中表现最优;肺癌中,纯文本模型最佳(F1=0.829)。深度学习模型始终优于传统模型。结肠癌因样本量最小,表现最差,凸显数据量重要性。SHAP分析显示不同癌种中各模态重要性差异显著。融合策略各有优劣,中间融合总体表现最佳,但应根据数据特征与实际需求选择。
原文摘要 · Abstract (English)
Multimodal Machine Learning offers a holistic view of a patient's status, integrating structured and unstructured data from electronic health records (EHR). We propose a framework to predict metastasis risk one month prior to diagnosis, using six months of clinical history from EHR data. Data from four cancer cohorts collected at Karolinska University Hospital (Stockholm, Sweden) were analyzed: breast (n = 743), colon (n = 387), lung (n = 870), and prostate (n = 1890). The dataset included demographics, comorbidities, laboratory results, medications, and clinical text. We compared traditional and deep learning classifiers across single modalities and multimodal combinations, using various fusion strategies and a Transparent Reporting of a multivariable prediction model for Individual Prognosis Or Diagnosis (TRIPOD) 2a design, with an 80-20 development-validation split to ensure a rigorous, repeatable evaluation. Performance was evaluated using AUROC, AUPRC, F1 score, sensitivity, and specificity. We then employed a multimodal adaptation of SHAP to analyze the classifiers' reasoning. Intermediate fusion achieved the highest F1 scores on breast (0.845), colon (0.786), and prostate cancer (0.845), demonstrating strong predictive performance. For lung cancer, the intermediate fusion achieved an F1 score of 0.819, while the text-only model achieved the highest, with an F1 score of 0.829. Deep learning classifiers consistently outperformed traditional models. Colon cancer, the smallest cohort, had the lowest performance, highlighting the importance of sufficient training data. SHAP analysis showed that the relative importance of modalities varied across cancer types. Fusion strategies offer distinct strengths and weaknesses. Intermediate fusion consistently delivered the best results, but strategy choices should align with data characteristics and organizational needs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。