用多维聚类识别免疫缺陷病特征,助力罕见病早期诊断。
A Multi-Dimensional Clustering Approach for Identifying Inborn Errors of Immunity

- 构建数据清洗与聚类分析流水线,从免疫检验数据提取疾病特征。
- 通过超参数调优实现模式识别,发现潜在新型免疫缺陷病模式。
- 为罕见病研究提供可复用的数据工具包,适合医疗数据科学家使用。
罕见病如先天性免疫缺陷(IEI)需尽早诊断以防止器官损伤并提升生活质量。然而,电子健康记录(EHR)数据的获取与整理困难,制约了基于数据的IEI趋势分析。当前在IEI领域用于模式识别的机器学习算法及系统化处理复杂医学数据的方法仍较匮乏。本文提出一套包含数据清洗与机器学习聚类算法的分析流程,旨在从国家级数据库中识别新型罕见病模式,并提取与IEI相关的特征。该方法将原始免疫学检验数据转化为向量表示,并结合超参数调优进行聚类分析,以实现疾病模式识别。本研究提升了对IEI特征的认知,开发了针对罕见病人群分析的数据工具包,并推动复杂医疗记录向无监督学习可解析的数据结构转化。
原文摘要 · Abstract (English)
Rare diseases such as inborn errors of immunity (IEI) require early diagnosis to prevent end organ damage and improve quality of life. Hurdles in accessing and curating large scale electronic health record (EHR) data limit routine data driven analyses to remain on the forefront of IEI and other rare disease trends. Development of machine learning (ML) algorithms in IEI for pattern recognition as well as published methodology examining how to systematically process and integrate complex medical data is limited. Our proposed pipeline, including data curation and ML clustering algorithms, is designed to recognize novel rare disease patterns and extract IEI- associated features from a national data registry. Our methodology for EHR data formatting and processing presents the pipeline that transforms raw immunologic lab data into vectors. This is further combined with hyperparameter tuning for diseases pattern recognition via clustering. This study refines IEI feature awareness, develops data tool kits for rare disease populations analysis, and expands on transforming complex medical records in data structures interpretable by unsupervised ML.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。