arXiv:2507.08189physics.med-phcs.LG2025-07被引 6

用少量标注数据+无标签数据,提升肺癌生存预测准确率与稳定性。

Robust Semi-Supervised CT Radiomics for Lung Cancer Prognosis: Cost-Effective Learning with Limited Labels and SHAP Interpretation

  • 采用半监督学习框架,利用无标签数据增强模型性能。
  • 仅10%标注数据时,生存预测准确率达0.90(交叉验证)和0.88(外部测试)。
  • 结合SHAP分析可解释预测结果,适合临床部署与医学AI研究者。

背景:CT影像对肺癌管理至关重要,为基于AI的预后分析提供详细可视化。然而,监督学习(SL)模型需大量标注数据,限制了在标注稀缺场景中的实际应用。方法:分析来自12个数据集的977例患者CT扫描,使用PyRadiomics提取1218个放射组学特征(基于拉普拉斯高斯和小波滤波器),通过56种特征选择与提取算法及27种分类器进行降维与建模。构建半监督学习(SSL)框架,结合478例无标签和499例有标签病例,测试三种情景下模型敏感性:逐步增加标注数据、扩大无标签数据、同时扩展两者(从10%到100%)。采用SHAP分析解释预测结果,并在两个外部队列中进行交叉验证与外部测试。结果:相比SL,SSL显著提升生存预测性能,最高提升达17%。最优模型(随机森林+XGBoost)在交叉验证中达到0.90准确率,在外部测试中达0.88。SHAP分析显示,两种方法均增强特征区分能力,尤其在生存期超过4年的患者中表现更优。当仅使用10%标注数据时,SSL仍具强性能,且外部测试中方差更低,体现其鲁棒性与成本效益。结论:提出一种低成本、稳定、可解释的半监督学习框架,通过融合无标签数据与SHAP可解释性,显著提升肺癌生存预测的性能、泛化性与临床适用性。

原文摘要 · Abstract (English)

Background: CT imaging is vital for lung cancer management, offering detailed visualization for AI-based prognosis. However, supervised learning SL models require large labeled datasets, limiting their real-world application in settings with scarce annotations. Methods: We analyzed CT scans from 977 patients across 12 datasets extracting 1218 radiomics features using Laplacian of Gaussian and wavelet filters via PyRadiomics Dimensionality reduction was applied with 56 feature selection and extraction algorithms and 27 classifiers were benchmarked A semi supervised learning SSL framework with pseudo labeling utilized 478 unlabeled and 499 labeled cases Model sensitivity was tested in three scenarios varying labeled data in SL increasing unlabeled data in SSL and scaling both from 10 percent to 100 percent SHAP analysis was used to interpret predictions Cross validation and external testing in two cohorts were performed. Results: SSL outperformed SL, improving overall survival prediction by up to 17 percent. The top SSL model, Random Forest plus XGBoost classifier, achieved 0.90 accuracy in cross-validation and 0.88 externally. SHAP analysis revealed enhanced feature discriminability in both SSL and SL, especially for Class 1 survival greater than 4 years. SSL showed strong performance with only 10 percent labeled data, with more stable results compared to SL and lower variance across external testing, highlighting SSL's robustness and cost effectiveness. Conclusion: We introduced a cost-effective, stable, and interpretable SSL framework for CT-based survival prediction in lung cancer, improving performance, generalizability, and clinical readiness by integrating SHAP explainability and leveraging unlabeled data.

肺癌预后半监督学习放射组学SHAP解释

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。