用机器学习分析越南空气污染、生态退化与肺癌的关联
Application of machine learning models to predict the relationship between air pollution, ecosystem degradation, and health disparities and lung cancer in Vietnam
- 结合健康数据与环境指标,用机器学习挖掘肺癌风险模式
- 随机森林和SVM模型准确率达99%,揭示关键影响因素
- 适合关注环境健康与公共卫生政策的研究者参考
肺癌是全球主要死因之一,越南亦不例外。2020年越南肺癌新增病例26,262例,死亡病例23,797例,占所有癌症病例的14.4%,位居第二死因(仅次于肝癌)。随着越南疾病负担加重,肺癌持续引发高度关注。在气候变化、多种污染、森林砍伐及现代生活方式背景下,肺癌风险处于高危状态。为深入理解越南特有的社会经济与生态背景下肺癌成因,本研究整合患者健康记录与环境指标(包括森林砍伐率、绿地覆盖率、空气污染水平及肺癌风险)等大数据,来自官方公开平台。通过热力图、信息增益、p值、斯皮尔曼相关性分析,识别潜在因果关联;并应用决策树、随机森林、支持向量机(SVM)、K均值聚类等机器学习模型挖掘癌症风险模式。实验结果显示,随机森林、SVM与主成分分析(PCA)表现优异,准确率达99%;而K均值聚类准确率仅10%,不适用于该数据集。
原文摘要 · Abstract (English)
Lung cancer is one of the major causes of death worldwide, and Vietnam is not an exception. This disease is the second most common type of cancer globally and the second most common cause of death in Vietnam, just after liver cancer, with 23,797 fatal cases and 26,262 new cases, or 14.4% of the disease in 2020. Recently, with rising disease rates in Vietnam causing a huge public health burden, lung cancer continues to hold the top position in attention and care. Especially together with climate change, under a variety of types of pollution, deforestation, and modern lifestyles, lung cancer risks are on red alert, particularly in Vietnam. To understand more about the severe disease sources in Vietnam from a diversity of key factors, including environmental features and the current health state, with a particular emphasis on Vietnam's distinct socioeconomic and ecological context, we utilize large datasets such as patient health records and environmental indicators containing necessary information, such as deforestation rate, green cover rate, air pollution, and lung cancer risks, that is collected from well-known governmental sharing websites. Then, we process and connect them and apply analytical methods (heatmap, information gain, p-value, spearman correlation) to determine causal correlations influencing lung cancer risks. Moreover, we deploy machine learning (ML) models (Decision Tree, Random Forest, Support Vector Machine, K-mean clustering) to discover cancer risk patterns. Our experimental results, leveraged by the aforementioned ML models to identify the disease patterns, are promising, particularly, the models as Random Forest, SVM, and PCA are working well on the datasets and give high accuracy (99%), however, the K means clustering has very low accuracy (10%) and does not fit the datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。