用机器学习预测乳腺癌患者是否需化疗或激素治疗,准确率达77%。
A Machine Learning Framework for Breast Cancer Treatment Classification Using a Novel Dataset
- 基于TCGA数据集,用梯度提升模型预测治疗方案。
- 最高准确率77.18%,AUROC达0.8252,性能稳定。
- 模型可解释性强,适合临床辅助决策参考。
乳腺癌(BC)仍是全球重大健康挑战,其分子与临床异质性使个体化治疗选择复杂化。本研究利用癌症基因组图谱(TCGA)乳腺癌临床数据集,构建机器学习模型以预测患者接受化疗或激素治疗的可能性。模型采用五折交叉验证训练,并通过准确率、精确率、召回率、特异性、敏感性、F1分数及受试者工作特征曲线下面积(AUROC)评估性能。通过自助法评估模型不确定性,借助SHAP值提升可解释性。在测试模型中,梯度提升机(GBM)表现最优(准确率=0.7718,AUROC=0.8252),其次为极端梯度提升(XGBoost)(准确率=0.7557,AUROC=0.8044)和自适应提升(AdaBoost)(准确率=0.7552,AUROC=0.8016)。结果表明,机器学习可通过数据驱动洞察支持个性化乳腺癌治疗决策。
原文摘要 · Abstract (English)
Breast cancer (BC) remains a significant global health challenge, with personalized treatment selection complicated by the disease's molecular and clinical heterogeneity. BC treatment decisions rely on various patient-specific clinical factors, and machine learning (ML) offers a powerful approach to predicting treatment outcomes. This study utilizes The Cancer Genome Atlas (TCGA) breast cancer clinical dataset to develop ML models for predicting the likelihood of undergoing chemotherapy or hormonal therapy. The models are trained using five-fold cross-validation and evaluated through performance metrics, including accuracy, precision, recall, specificity, sensitivity, F1-score, and area under the receiver operating characteristic curve (AUROC). Model uncertainty is assessed using bootstrap techniques, while SHAP values enhance interpretability by identifying key predictors. Among the tested models, the Gradient Boosting Machine (GBM) achieves the highest stable performance (accuracy = 0.7718, AUROC = 0.8252), followed by Extreme Gradient Boosting (XGBoost) (accuracy = 0.7557, AUROC = 0.8044) and Adaptive Boosting (AdaBoost) (accuracy = 0.7552, AUROC = 0.8016). These findings underscore the potential of ML in supporting personalized breast cancer treatment decisions through data-driven insights.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。