arXiv:2410.17494eess.IVcs.CV2024-10被引 5

融合影像与非影像数据,提升疾病分类准确率与可解释性

Enhancing Multimodal Medical Image Classification using Cross-Graph Modal Contrastive Learning

  • 构建跨模态图结构,通过对比学习对齐多模态特征
  • 在帕金森病和黑色素瘤数据集上准确率显著优于单一模态方法
  • 适合需要多源数据融合的医学诊断场景

医学图像分类是疾病诊断的关键环节,常依赖深度学习技术。然而,传统方法多聚焦于单一模态的医学图像数据,忽视了非图像患者数据的整合。本文提出一种跨图模态对比学习(CGMCL)框架,用于处理来自不同数据域的多模态结构化数据,以提升医学图像分类性能。该模型通过构建跨模态图结构,并利用对比学习将多模态特征对齐至共享隐空间,有效融合图像与非图像数据。引入模态间特征缩放模块,进一步缩小异质模态间的表示差距,优化表征学习过程。在帕金森病(PD)数据集和公开黑色素瘤数据集上进行评估,结果表明CGMCL在分类准确率、可解释性及早期疾病预测方面均优于传统单模态方法,且在多类黑色素瘤分类中表现更优。该框架为医学图像分类提供了新思路,提升了疾病可解释性与预测能力。

原文摘要 · Abstract (English)

The classification of medical images is a pivotal aspect of disease diagnosis, often enhanced by deep learning techniques. However, traditional approaches typically focus on unimodal medical image data, neglecting the integration of diverse non-image patient data. This paper proposes a novel Cross-Graph Modal Contrastive Learning (CGMCL) framework for multimodal structured data from different data domains to improve medical image classification. The model effectively integrates both image and non-image data by constructing cross-modality graphs and leveraging contrastive learning to align multimodal features in a shared latent space. An inter-modality feature scaling module further optimizes the representation learning process by reducing the gap between heterogeneous modalities. The proposed approach is evaluated on two datasets: a Parkinson's disease (PD) dataset and a public melanoma dataset. Results demonstrate that CGMCL outperforms conventional unimodal methods in accuracy, interpretability, and early disease prediction. Additionally, the method shows superior performance in multi-class melanoma classification. The CGMCL framework provides valuable insights into medical image classification while offering improved disease interpretability and predictive capabilities.

多模态学习医学图像对比学习疾病预测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。