arXiv:2505.22592eess.IVcs.CV2025-05

对比两种模型在3D CT肺癌突变与分期检测中的表现

Comparative Analysis of Machine Learning Models for Lung Cancer Mutation Detection and Staging Using 3D CT Scans

  • 用领域预训练+XGBoost做突变检测,自监督+注意力多实例学习做分期
  • 突变检测准确率最高达88.3%,分期预测准确率达79.7%
  • 适合关注肺癌精准诊疗的临床医生和医学影像研究者

肺癌是全球癌症死亡的主要原因,非侵入性检测关键突变和分期对改善患者预后至关重要。本文对比了两种机器学习模型——基于领域特定预训练的监督模型FMCIB+XGBoost,以及基于自监督学习与注意力机制多实例学习的Dinov2+ABMIL——在斯坦福放射基因组学与Lung-CT-PT-Dx数据集上的表现。在KRAS和EGFR突变检测任务中,FMCIB+XGBoost表现更优,准确率分别为0.846和0.883;在癌症分期任务中,Dinov2+ABMIL展现出良好泛化能力,在Lung-CT-PT-Dx数据集上T期预测准确率达0.797,表明自监督学习在跨数据集适应中的潜力。结果凸显监督模型在突变检测中的临床价值,并提示自监督学习在分期泛化方面的前景,同时指出突变敏感性仍有提升空间。

原文摘要 · Abstract (English)

Lung cancer is the leading cause of cancer mortality worldwide, and non-invasive methods for detecting key mutations and staging are essential for improving patient outcomes. Here, we compare the performance of two machine learning models - FMCIB+XGBoost, a supervised model with domain-specific pretraining, and Dinov2+ABMIL, a self-supervised model with attention-based multiple-instance learning - on 3D lung nodule data from the Stanford Radiogenomics and Lung-CT-PT-Dx cohorts. In the task of KRAS and EGFR mutation detection, FMCIB+XGBoost consistently outperformed Dinov2+ABMIL, achieving accuracies of 0.846 and 0.883 for KRAS and EGFR mutations, respectively. In cancer staging, Dinov2+ABMIL demonstrated competitive generalization, achieving an accuracy of 0.797 for T-stage prediction in the Lung-CT-PT-Dx cohort, suggesting SSL's adaptability across diverse datasets. Our results emphasize the clinical utility of supervised models in mutation detection and highlight the potential of SSL to improve staging generalization, while identifying areas for enhancement in mutation sensitivity.

肺癌3D CT突变检测自监督

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。