自动构建多源病历数据的层级图谱,解决医疗代码异构难题
Automated Hierarchical Graph Construction for Multi-source Electronic Health Records
- 用神经最优传输对齐多机构医疗代码,学习双曲嵌入构建层级图
- 整合语言模型与共现模式,提升医学概念语义关系捕捉能力
- 首次实现无结构实验室代码的自动化层级生成,适合临床研究者
电子健康记录(EHR)包含诊断、药物、检验等多样临床数据,对转化研究具有重要价值。跨机构融合EHR可支持大规模、可泛化的研究,揭示罕见病和人群多样性,但受限于医学编码异质性、机构术语差异及缺乏标准数据结构。这些障碍影响分析的可解释性、可比性和可扩展性,亟需稳健方法实现分布式异构数据的标准化与洞察提取。为此,我们提出MASH(多源自动结构化层级框架),通过神经最优传输对齐多机构医疗代码,并利用学习的双曲嵌入构建层级图。训练中融合预训练语言模型、共现模式、文本描述与监督标签,更有效地捕获医学概念间的语义与层级关系。在真实世界EHR数据(含诊断、药物、检验编码)上的应用表明,MASH生成了可解释的层级图,促进异构临床数据的理解与导航。尤为关键的是,它首次实现了对非结构化本地检验代码的自动化层级构建,为下游应用建立了基础参考。
原文摘要 · Abstract (English)
Electronic Health Records (EHRs), comprising diverse clinical data such as diagnoses, medications, and laboratory results, hold great promise for translational research. EHR-derived data have advanced disease prevention, improved clinical trial recruitment, and generated real-world evidence. Synthesizing EHRs across institutions enables large-scale, generalizable studies that capture rare diseases and population diversity, but remains hindered by the heterogeneity of medical codes, institution-specific terminologies, and the absence of standardized data structures. These barriers limit the interpretability, comparability, and scalability of EHR-based analyses, underscoring the need for robust methods to harmonize and extract meaningful insights from distributed, heterogeneous data. To address this, we propose MASH (Multi-source Automated Structured Hierarchy), a fully automated framework that aligns medical codes across institutions using neural optimal transport and constructs hierarchical graphs with learned hyperbolic embeddings. During training, MASH integrates information from pre-trained language models, co-occurrence patterns, textual descriptions, and supervised labels to capture semantic and hierarchical relationships among medical concepts more effectively. Applied to real-world EHR data, including diagnosis, medication, and laboratory codes, MASH produces interpretable hierarchical graphs that facilitate the navigation and understanding of heterogeneous clinical data. Notably, it generates the first automated hierarchies for unstructured local laboratory codes, establishing foundational references for downstream applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。