arXiv:2505.22554stat.MLcs.LG2025-05被引 1

用尾部相关性筛选糖尿病风险特征,更关注极端高危人群的预测因子。

A Copula Based Supervised Filter for Feature Selection in Diabetes Risk Prediction Using Machine Learning

  • 基于Gumbel-copula设计上尾一致评分,捕捉特征与高风险同时极端的能力。
  • 在CDC数据集上从21个特征精简至10个,提升模型效率且保持性能稳定。
  • 适合关注高危人群的公共卫生与临床风险预测场景,结果可解释性强。

在医学预测建模中,有效特征选择对模型稳健性和可解释性至关重要,尤其当风险因素在极端患者群体中更为关键时。许多标准筛选方法侧重平均关联,可能遗漏集中在分布尾部的相关预测因子。本文提出一种基于Gumbel-copula的高效监督过滤器,使用上尾一致性得分(lambda U),该得分是Kendall's tau的单调变换,用于按特征与正类同时极端的可能性排序。我们在两个糖尿病数据集上对比了四种基线方法(互信息、mRMR、ReliefF、L1/Elastic-Net):大规模公共健康调查(CDC,N=253,680)和临床基准(PIMA,N=768)。分析包括统计检验、置换重要性及鲁棒性检查。在CDC数据集上,所提方法最快,将21个特征缩减至10个(约52%),在使用全部特征的基础上实现微小但显著的性能权衡,优于传统滤波器(互信息、mRMR),并可与强基线ReliefF相当。在PIMA数据集(8个预测因子)上,所得排名达到最高数值ROC-AUC,但成对DeLong检验未显示显著差异;因此,在低维设置下作为仅排名的合理性验证。跨两个数据集,lambda U筛选器均识别出具有临床合理性的预测因子,提供了一种高效、可解释的预筛步骤,可补充标准特征选择方法在公共卫生与临床风险预测中的应用。

原文摘要 · Abstract (English)

Effective feature selection is critical for robust and interpretable predictive modeling in medicine, especially when risk factors matter most in extreme patient strata. Many standard selectors emphasize average associations and can miss predictors whose relevance is concentrated in the distribution tails. We propose a computationally efficient supervised filter based on a Gumbel-copula implied upper-tail concordance score (lambda U), defined as a monotone transformation of Kendall's tau, to rank features by their tendency to be simultaneously extreme with the positive class. We compare against four common baselines (Mutual Information, mRMR, ReliefF, and L1/Elastic-Net) across four classifiers on two diabetes datasets: a large-scale public health survey (CDC, N=253,680) and a clinical benchmark (PIMA, N=768). Analyses include statistical testing, permutation importance, and robustness checks. On CDC, the proposed selector is the fastest and reduces 21 features to 10 (approx 52%). This yields a small but statistically significant trade-off relative to using all features, while performing better than standard filters (Mutual Information, mRMR) and comparably to the strong ReliefF baseline. On PIMA (8 predictors), the resulting ranking attains the highest ROC-AUC numerically, though paired DeLong tests show no significant differences versus strong baselines; PIMA therefore serves as a ranking-only sanity check in a low-dimensional setting. Across both datasets, the lambda U-based selector highlights clinically coherent predictors and provides an efficient, interpretable screening step that can complement standard feature-selection methods in public health and clinical risk prediction.

特征选择糖尿病预测尾部相关可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。