用可解释机器学习,在低资源地区精准筛查早期肾病。
Community-Based Early-Stage Chronic Kidney Disease Screening using Explainable Machine Learning for Low-Resource Settings
- 按临床意义分组变量,结合十种特征选择法找关键风险因子。
- 模型平衡准确率达90.4%,在多国数据集上敏感性达78%-98%。
- 只需少量易得指标,适合资源有限地区的社区筛查使用。
早期发现慢性肾病(CKD)对防止进展至终末期肾病至关重要。然而,现有筛查工具主要基于高收入国家人群开发,在孟加拉国及南亚地区表现不佳,因当地风险特征不同。多数工具依赖简单加法评分,且基于晚期肾病患者数据,难以捕捉风险因素间复杂交互,预测早期肾病能力有限。本研究旨在为孟加拉国及南亚低资源环境开发并评估一种可解释的机器学习(ML)框架,用于社区早期肾病筛查。基于孟加拉国的社区数据集构建预测模型,将变量按临床意义分组,采用十种互补的特征选择方法识别稳健预测子集,并评估十二种机器学习分类器,使用嵌套交叉验证。模型性能与现有筛查工具对比,并在印度、阿联酋和孟加拉国三个独立数据集上进行外部验证。采用SHAP解释模型预测。基于RFECV选择的特征子集训练的模型达到90.40%的平衡准确率;仅包含最少非病理学检测指标的模型也表现出89.23%的平衡准确率,常优于更大或完整特征集。相比现有工具,新模型显著提升准确率与灵敏度,且所需输入更少、更易获取。外部验证证实其强泛化能力,灵敏度为78%至98%。
原文摘要 · Abstract (English)
Early detection of chronic kidney disease (CKD) is essential for preventing progression to end-stage renal disease. However, existing screening tools - primarily developed using populations from high-income countries - often underperform in Bangladesh and South Asia, where risk profiles differ. Most of these tools rely on simple additive scoring functions and are based on data from patients with advanced-stage CKD. Consequently, they fail to capture complex interactions among risk factors and are limited in predicting early-stage CKD. Our objective was to develop and evaluate an explainable machine learning (ML) framework for community-based early-stage CKD screening for low-resource settings, tailored to the Bangladeshi and South Asian population context. A community-based CKD dataset from Bangladesh was used to develop predictive models. Variables were organized into clinically meaningful feature groups, and ten complementary feature selection methods were applied to identify robust predictor subsets. Twelve ML classifiers were evaluated using nested cross-validation. Model performance was benchmarked against established CKD screening tools and externally validated on three independent datasets from India, the UAE, and Bangladesh. SHAP was used to interpret model predictions. An ML model trained on an RFECV-selected feature subset achieved a balanced accuracy of 90.40%, whereas minimal non-pathology-test features demonstrated excellent predictive capability with a balanced accuracy of 89.23%, often outperforming larger or full feature sets. Compared with existing screening tools, the proposed models achieved substantially higher accuracy and sensitivity while requiring fewer and more accessible inputs. External validation confirmed strong generalizability with 78% to 98% sensitivity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。