用癌症数据库建模预测肝癌肺转移,帮医生提前识别高危患者。
Prediction of Lung Metastasis from Hepatocellular Carcinoma using the SEER Database
- 基于SEER数据构建机器学习模型,融合临床与人口统计信息。
- 随机森林和MLP模型表现最佳,AUROC达0.82,但精确率偏低。
- 引入召回优化损失函数和集成策略,提升对高危患者的识别能力。
肝细胞癌(HCC)是导致癌症死亡的主要原因之一,肺转移是最常见的远处转移部位,显著恶化预后。尽管临床与人口统计数据日益丰富,针对HCC肺转移的预测模型仍存在范围有限、临床适用性不足的问题。本研究利用癌症监测、流行病学与结果数据库(SEER)数据,构建并验证了一个端到端的机器学习流程。评估了随机森林、XGBoost、逻辑回归及多层感知机(MLP)神经网络三种模型。各模型表现出较高的AUROC与召回率,其中随机森林与MLP表现最优(AUROC = 0.82)。然而,所有模型的精确率均较低,表明准确识别阳性病例仍具挑战。为此,我们设计了一种结合召回优化的自定义损失函数,使MLP模型实现最高敏感性。集成方法进一步通过融合各模型优势提升了整体召回率。特征重要性分析揭示手术状态、肿瘤分期和随访时长为关键预测因子,凸显临床干预与疾病进展在转移预测中的作用。尽管研究展示了机器学习在识别高风险患者方面的潜力,但仍受限于数据不平衡、特征标注不全以及预测精确率低等问题。未来工作应利用扩大的SEER数据集,改进数据填补技术,并探索先进预训练模型以提升预测准确性与临床实用性。
原文摘要 · Abstract (English)
Hepatocellular carcinoma (HCC) is a leading cause of cancer-related mortality, with lung metastases being the most common site of distant spread and significantly worsening prognosis. Despite the growing availability of clinical and demographic data, predictive models for lung metastasis in HCC remain limited in scope and clinical applicability. In this study, we develop and validate an end-to-end machine learning pipeline using data from the Surveillance, Epidemiology, and End Results (SEER) database. We evaluated three machine learning models (Random Forest, XGBoost, and Logistic Regression) alongside a multilayer perceptron (MLP) neural network. Our models achieved high AUROC values and recall, with the Random Forest and MLP models demonstrating the best overall performance (AUROC = 0.82). However, the low precision across models highlights the challenges of accurately predicting positive cases. To address these limitations, we developed a custom loss function incorporating recall optimization, enabling the MLP model to achieve the highest sensitivity. An ensemble approach further improved overall recall by leveraging the strengths of individual models. Feature importance analysis revealed key predictors such as surgery status, tumor staging, and follow up duration, emphasizing the relevance of clinical interventions and disease progression in metastasis prediction. While this study demonstrates the potential of machine learning for identifying high-risk patients, limitations include reliance on imbalanced datasets, incomplete feature annotations, and the low precision of predictions. Future work should leverage the expanding SEER dataset, improve data imputation techniques, and explore advanced pre-trained models to enhance predictive accuracy and clinical utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。