利用单次就诊记录和共现症状提升罕见事件建模效果
Salvaging Forbidden Treasure in Medical Data: Utilizing Surrogate Outcomes and Single Records for Rare Event Modeling
- 用共现症状作为替代指标,融合历史与单次就诊数据
- 单次就诊患者占自杀案例70%-80%,显著提升模型性能
- 适合医疗罕见事件建模,尤其对数据稀缺场景有效
电子健康记录(EHR)和医疗理赔数据中蕴藏着研究罕见但关键事件的巨大潜力,如自杀未遂。传统方法将自杀未遂视为单一结局,并排除仅有一条记录的患者,因其缺乏历史信息。然而,这些仅在一次就诊中被记录的患者,可能占所有自杀未遂病例的70%至80%。本文提出一种混合集成学习框架,利用共现结局作为代理指标,并挖掘单次记录数据中的宝贵信息。该方法通过有监督学习建模主结局(如自杀)与代理结局(如精神疾病)之间的潜在变量关系,同时通过无监督学习利用单次记录数据,共享潜在表示。基于康涅狄格州住院数据的实验证明,单次记录和共现诊断确实包含重要信息,整合后可显著提升自杀风险预测性能。
原文摘要 · Abstract (English)
The vast repositories of Electronic Health Records (EHR) and medical claims hold untapped potential for studying rare but critical events, such as suicide attempt. Conventional setups often model suicide attempt as a univariate outcome and also exclude any ``single-record'' patients with a single documented encounter due to a lack of historical information. However, patients who were diagnosed with suicide attempts at the only encounter could, to some surprise, represent a substantial proportion of all attempt cases in the data, as high as 70--80%. We innovate a hybrid and integrative learning framework to leverage concurrent outcomes as surrogates and harness the forbidden yet precious information from single-record data. Our approach employs a supervised learning component to learn the latent variables that connect primary (e.g., suicide) and surrogate outcomes (e.g., mental disorders) to historical information. It simultaneously employs an unsupervised learning component to utilize the single-record data, through the shared latent variables. As such, our approach offers a general strategy for information integration that is crucial to modeling rare conditions and events. With hospital inpatient data from Connecticut, we demonstrate that single-record data and concurrent diagnoses indeed carry valuable information, and utilizing them can substantially improve suicide risk modeling.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。