用多模态电子病历数据挖掘疾病细分类型,提升个性化诊疗预测能力。
MixEHR-Nest: Identifying Subphenotypes within Electronic Health Records through Hierarchical Guided-Topic Modeling
- 基于专家标注的病种编码引导层次化主题模型,挖掘疾病内部子类型。
- 在3.8万例重症患者和130万例普通患者数据中验证,精准识别出糖尿病亚型及预后差异。
- 适用于临床研究、疾病分型与长期随访分析,适合医疗数据科学家使用。
从电子健康记录(EHR)中自动识别疾病亚表型,有助于理解异质性疾病并推动个性化医疗。现有方法或针对特定疾病以提升可解释性,或生成粗粒度表型主题而忽略细微疾病模式。本研究提出一种引导式主题模型 MixEHR-Nest,从数千种疾病中利用多模态 EHR 数据推断亚表型主题。该模型在每个表型主题下检测多个子主题,其先验由专家构建的表型概念(如 PheCodes、CCS 码)引导。我们在两个 EHR 数据集上评估:(1) MIMIC-III,包含波士顿贝斯以色列女执事医学中心(BIDMC)38,000 名重症监护室(ICU)患者;(2) PopHR,涵盖蒙特利尔130万例患者的医疗行政数据库。实验表明,MixEHR-Nest 能识别出各表型内具有显著特征的亚表型,且对疾病进展与严重程度具有预测力。例如,它通过 CCS 码成功区分了1型与2型糖尿病,而原始代码无法区分两者。此外,该模型提升了重症患者短期死亡率及糖尿病患者初始胰岛素治疗的预测准确率,并揭示了亚表型的贡献。纵向分析发现,哮喘、白血病、癫痫和抑郁症等相同表型下存在不同年龄分布的亚表型。MixEHR-Nest 软件已开源:https://github.com/li-lab-mcgill/MixEHR-Nest。
原文摘要 · Abstract (English)
Automatic subphenotyping from electronic health records (EHRs)provides numerous opportunities to understand diseases with unique subgroups and enhance personalized medicine for patients. However, existing machine learning algorithms either focus on specific diseases for better interpretability or produce coarse-grained phenotype topics without considering nuanced disease patterns. In this study, we propose a guided topic model, MixEHR-Nest, to infer sub-phenotype topics from thousands of disease using multi-modal EHR data. Specifically, MixEHR-Nest detects multiple subtopics from each phenotype topic, whose prior is guided by the expert-curated phenotype concepts such as Phenotype Codes (PheCodes) or Clinical Classification Software (CCS) codes. We evaluated MixEHR-Nest on two EHR datasets: (1) the MIMIC-III dataset consisting of over 38 thousand patients from intensive care unit (ICU) from Beth Israel Deaconess Medical Center (BIDMC) in Boston, USA; (2) the healthcare administrative database PopHR, comprising 1.3 million patients from Montreal, Canada. Experimental results demonstrate that MixEHR-Nest can identify subphenotypes with distinct patterns within each phenotype, which are predictive for disease progression and severity. Consequently, MixEHR-Nest distinguishes between type 1 and type 2 diabetes by inferring subphenotypes using CCS codes, which do not differentiate these two subtype concepts. Additionally, MixEHR-Nest not only improved the prediction accuracy of short-term mortality of ICU patients and initial insulin treatment in diabetic patients but also revealed the contributions of subphenotypes. For longitudinal analysis, MixEHR-Nest identified subphenotypes of distinct age prevalence under the same phenotypes, such as asthma, leukemia, epilepsy, and depression. The MixEHR-Nest software is available at GitHub: https://github.com/li-lab-mcgill/MixEHR-Nest.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。