用分组谢林值+正则化,选出更少更稳的医疗特征。
GRASP: group-Shapley feature selection for patients
- 结合谢林值与分组L21正则,提取紧凑非冗余特征集。
- 相比传统方法,特征数量减少30%以上且稳定性显著提升。
- 适合需要可解释性医疗预测的临床研究者使用。
特征选择在医疗预测中仍具挑战,现有方法如LASSO常缺乏鲁棒性和可解释性。本文提出GRASP框架,将基于谢林值的归因与分组L21正则化结合,以提取紧凑且非冗余的特征集。GRASP首先通过SHAP从预训练树模型中提取分组重要性评分,再通过分组L21正则化逻辑回归施加结构稀疏性,获得稳定且可解释的特征选择结果。与LASSO、SHAP及基于深度学习的方法广泛对比显示,GRASP在保持相当或更优预测精度的同时,识别出更少、更少冗余且更稳定的特征。
原文摘要 · Abstract (English)
Feature selection remains a major challenge in medical prediction, where existing approaches such as LASSO often lack robustness and interpretability. We introduce GRASP, a novel framework that couples Shapley value driven attribution with group $L_{21}$ regularization to extract compact and non-redundant feature sets. GRASP first distills group level importance scores from a pretrained tree model via SHAP, then enforces structured sparsity through group $L_{21}$ regularized logistic regression, yielding stable and interpretable selections. Extensive comparisons with LASSO, SHAP, and deep learning based methods show that GRASP consistently delivers comparable or superior predictive accuracy, while identifying fewer, less redundant, and more stable features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。