PEHRT统一整合医疗数据,让跨机构研究更便捷。
A Common Pipeline for Harmonizing Electronic Health Record Data for Translational Research
- 构建通用流程,用标准化术语统一异构医疗数据
- 借助预训练模型生成语义嵌入,解决本地编码不一致问题
- 支持多机构协作训练,保护隐私同时提升模型泛化能力
尽管电子健康记录(EHR)数据日益丰富,研究人员在开展转化研究时仍面临数据复杂、异构性强及缺乏标准化工具与文档的挑战。为此,我们提出PEHRT——一个面向转化研究的EHR数据统一处理通用管道。该管道包含开源代码、可视化工具和详细文档,可直接使用。它通过将结构化与非结构化EHR数据映射至标准术语集,实现跨系统的一致性。对于未映射或本地异构编码,PEHRT利用表示学习与预训练语言模型生成稳健嵌入,捕捉多中心间的语义关联,缓解异质性,支持集成分析。该框架还支持跨机构联合训练,通过共享表示实现模型协同优化,无需交换个体数据。其数据模型无关性使其可无缝部署于各类医疗系统,生成可互操作、研究就绪的数据集。通过降低技术门槛,PEHRT助力研究者将原始临床数据转化为可复现、可分析的科研资源。
原文摘要 · Abstract (English)
Despite the growing availability of Electronic Health Record (EHR) data, researchers often face substantial barriers in effectively using these data for translational research due to their complexity, heterogeneity, and lack of standardized tools and documentation. To address this critical gap, we introduce PEHRT, a common pipeline for harmonizing EHR data for translational research. PEHRT is a comprehensive, ready-to-use resource that includes open-source code, visualization tools, and detailed documentation to streamline the process of preparing EHR data for analysis. The pipeline provides tools to harmonize structured and unstructured EHR data to standardized ontologies to ensure consistency across diverse coding systems. In the presence of unmapped or heterogeneous local codes, PEHRT further leverages representation learning and pre-trained language models to generate robust embeddings that capture semantic relationships across sites to mitigate heterogeneity and enable integrative downstream analyses. PEHRT also supports cross-institutional co-training through shared representations, allowing participating sites to collaboratively refine embeddings and enhance generalizability without sharing individual-level data. The framework is data model-agnostic and can be seamlessly deployed across diverse healthcare systems to produce interoperable, research-ready datasets. By lowering the technical barriers to EHR-based research, PEHRT empowers investigators to transform raw clinical data into reproducible, analysis-ready resources for discovery and innovation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。