arXiv:2501.11535cs.CV2025-01被引 1

融合影像与临床数据,提升肝癌分期预测准确率

A baseline for machine-learning-based hepatocellular carcinoma diagnosis using multi-modal clinical data

  • 用图像与实验室数据联合建模,提取多模态特征
  • 模型对肝癌TNM分期预测准确率达89%,AUC达0.93
  • 证明仅靠单一数据源无法实现高精度诊断

本文旨在为肝细胞癌(HCC)的多模态数据分类提供基准,基于一个全新的开放多模态数据集,包含增强CT和MRI影像数据,以及临床检验数据与病例报告表。研究将向量化预处理后的表格数据特征与来自增强CT和MRI的放射组学特征进行整合。通过互信息法进行特征选择,采用XGBoost分类器预测TNM分期,结果达到准确率 $0.89 /pm 0.05$,AUC为 $0.93 /pm 0.03$。实验表明,只有结合影像与临床检验数据才能获得如此高的预测性能,验证了多模态分类在肝癌精准分期中的必要性。

原文摘要 · Abstract (English)

The objective of this paper is to provide a baseline for performing multi-modal data classification on a novel open multimodal dataset of hepatocellular carcinoma (HCC), which includes both image data (contrast-enhanced CT and MRI images) and tabular data (the clinical laboratory test data as well as case report forms). TNM staging is the classification task. Features from the vectorized preprocessed tabular data and radiomics features from contrast-enhanced CT and MRI images are collected. Feature selection is performed based on mutual information. An XGBoost classifier predicts the TNM staging and it shows a prediction accuracy of $0.89 \pm 0.05$ and an AUC of $0.93 \pm 0.03$. The classifier shows that this high level of prediction accuracy can only be obtained by combining image and clinical laboratory data and therefore is a good example case where multi-model classification is mandatory to achieve accurate results.

肝癌诊断多模态学习XGBoostTNM分期

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。