比较五种特征选择方法,提升阿片类药物滥用预测的准确性和可解释性。
A Comparative Study of Feature Selection Methods for EHR Diagnosis Codes in Opioid Use Disorder Prediction

- 基于梯度敏感性、树模型和大模型语义分析,系统对比五类特征选择方法。
- NTK敏感性方法在准确率与稳定性上表现最佳,支持适度特征规模下的高效建模。
- 大模型引导的筛选虽单独效果一般,但能补充临床有意义的诊断信号。
电子健康记录(EHR)中的预测建模面临高维、稀疏、噪声多且冗余的输入变量问题。大规模特征集不仅增加计算负担和过拟合风险,也使模型难以解释,限制了其在临床中的应用。本研究聚焦于诊断相关特征,对比五种特征选择方法在阿片类药物滥用(OUD)预测中的表现:复发富集、基于神经切线核(NTK)的早期梯度敏感性、LightGBM-SHAP、弹性网络(Elastic Net)以及大语言模型(LLM)引导的语义选择。采用统一的预处理与评估框架,从下游预测性能、重抽样稳定性及对罕见诊断码的表示能力三方面进行评估。结果表明,随着特征预算增大,性能持续提升但呈现边际递减;NTK敏感性方法在准确率与稳定性之间取得最佳平衡;而LLM引导的选择虽独立表现较低,却能提供互补的临床可解释信号。
原文摘要 · Abstract (English)
Feature selection is a critical step in electronic health record (EHR)-based predictive modeling, where input variables are often high-dimensional, sparse, noisy, and redundant. Large feature sets not only increase computational burden and overfitting risk, but also make model interpretation difficult, leading to limited usefulness in clinical settings. In this study, we focus on diagnosis-related features and compare five feature selection paradigms for opioid use disorder (OUD) prediction: recurrence enrichment, NTK-motivated early gradient sensitivity, LightGBM-SHAP, Elastic Net, and large language model (LLM)-guided semantic selection. We use a unified preprocessing and evaluation framework and assess each method by downstream predictive performance, resampling stability, and representation of infrequent diagnosis codes. Our results demonstrate that performance improves with larger feature budgets with diminishing returns beyond a moderate size. NTK sensitivity provides the best overall balance of accuracy and stability, and LLM-guided selection contributes complementary clinically meaningful signals despite lower standalone performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。