用全国电子病历数据建模,提升慢性鼻窦炎早期预测准确率。
Nationwide EHR-Based Chronic Rhinosinusitis Prediction Using Demographic-Stratified Models

- 分人群构建子模型,结合统计与模型重要性筛选特征。
- 整体AUC达0.8461,比最优基线提升0.0168。
- 适合关注慢病早筛与医疗资源优先分配的研究者。
慢性鼻窦炎(CRS)是一种常见且异质性强的炎症性疾病,导致显著健康负担和医疗成本。由于症状常与过敏性鼻炎等常见病重叠,且表型多样,早期识别困难。以往研究多依赖单机构队列,影响泛化性。本文利用来自「所有人」研究计划的全国纵向电子病历数据,基于诊断前两年的记录预测CRS。针对编码数据中存在的极端稀疏性和高维问题,提出混合特征筛选流程,将约11万候选代码压缩至100个可解释特征。为捕捉人口学差异,针对六类成人性别与生命阶段亚组分别训练分层模型,并进行子组特异性超参数调优。整体AUC达0.8461,较最佳基线提升0.0168。结果表明,常规采集的EHR数据可支持具有代表性的CRS风险分层,有助于初级诊疗中更早的分诊与转诊优先级判定。
原文摘要 · Abstract (English)
Chronic rhinosinusitis (CRS) is a common heterogeneous inflammatory disorder that causes substantial morbidity and healthcare costs. CRS is difficult to identify early from routine encounters, as symptom presentations overlap with common conditions such as allergic rhinitis, and heterogeneous phenotypes further obscure risk patterns. Prior predictive studies often rely on single-institutional cohorts , which reduce population-level generalizability. To overcome this, we leveraged nationwide longitudinal EHR data from the \textit{All of Us} Research Program to predict CRS diagnosis using two years of pre-diagnostic history. To address extreme feature sparsity and dimensionality in coded EHR data, we implemented a hybrid feature-selection pipeline that combines prevalence-based statistical screening with model-based importance ranking, compressing approximately 110,000 candidate codes into 100 interpretable features. To capture demographic heterogeneity, we trained demographic stratified models across six adult sex and life-stage subgroups with subgroup-specific hyperparameter tuning. Our framework achieved an overall AUC of 0.8461, improving discrimination by 0.0168 over the best baseline. These results demonstrate that routinely collected EHR data may support population-representative CRS risk stratification and inform earlier triage and referral prioritization in primary care.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。