arXiv:2607.11963cs.LGcs.AI2026-07综述

发现肾病预测模型性能虚高,主因是数据泄露和指标不可靠。

Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability

论文配图:Evaluating Reliability in Machine Learning Models for Early Chronic Kidney Disease Prediction: A Systematic Review of Data Leakage and Predictor Stability
图 1 · 摘自论文原文
  • 构建泄漏分类体系与量化评分框架,系统评估模型可靠性。
  • 有泄露研究平均准确率95.48%,无泄露仅80.2%,差距达15.28%。
  • 超80%预测因子不可复现,多数成果源于方法缺陷而非真能力。

利用机器学习进行慢性肾病早期检测在医疗计算领域备受关注。尽管进展迅速,但许多研究结果不一致且可能误导。主要问题包括数据泄露、缺乏时间序列患者记录以及临床指标报告不一致。本研究通过系统检索主要学术数据库,筛选出19篇相关研究,采用可解释机器学习方法进行综述。为评估方法可靠性,提出信息泄露的结构化分类体系与定量评分框架,系统分析各研究的可靠性。结果表明,泄露与性能虚高显著相关:高泄露研究平均准确率达95.48%,而无泄露研究仅为80.2%,提升约15.28%。此外,跨研究特征稳定性分析显示,仅有少数预测因子可重复,超过80%缺乏可靠性。总体而言,多数报告的性能提升源于方法学缺陷,而非真实预测能力。

原文摘要 · Abstract (English)

The early detection of Chronic Kidney Disease using machine learning has attracted significant interest in healthcare-related computer science. Despite rapid advancements in this field, many reported studies remain inconsistent and potentially misleading. A significant drawback is the lack of organized evaluation regarding methodological concerns. Key issues include data leakage, limited access to temporal patient records and inconsistency in reported clinical indicators. This research offers a systematic literature review of existing CKD prediction studies using interpretable machine learning techniques, where nineteen relevant studies were selected via systematic searches across major academic databases. To assess methodological reliability, this study introduces a structured taxonomy of information leakage and a quantitative leakage scoring framework to systematically evaluate reliability across CKD prediction studies. The analysis reveals a strong relationship between leakage and inflated performance. Here, High leakage-studies report an average accuracy of 95.48%, compared to 80.2% for leakage-free studies, reflecting an increase of approximately 15.28%. Furthermore, a cross-study feature stability analysis shows that only a small subset of predictors is consistently reproducible, with over 80% lacking reliability. Overall, the findings suggest that many reported performance improvements stem from methodological limitations rather than true predictive capability.

慢性肾病数据泄露模型评估可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。