arXiv:2502.08547cs.AI2025-02被引 8

构建跨机构医疗数据共享的统一语义空间,解决编码差异与隐私保护难题。

Representation learning to advance multi-institutional studies with electronic health record data from US and France

  • 基于图结构的表示学习框架,融合多源信息对齐异构医学术语
  • 在7个机构、双语言环境下实现跨系统临床模型稳定训练
  • 无需手动映射,兼顾隐私保护与数据一致性,适合多中心医疗研究

电子健康记录的广泛应用为转化临床研究带来新机遇,但受制于各机构间数据割裂及本地编码实践差异。尽管隐私保护协同学习可避免患者级数据共享,却无法解决临床概念表达不一致问题。本文提出一种基于图的框架,将数据标准化转化为可扩展的表示学习任务。该框架整合机构特定的汇总统计、经编纂的生物医学知识图谱及大语言模型提取的语义信息,共同学习一个共享语义空间,实现多样本、站点特异性术语的对齐,同时保障患者隐私。在7个机构和两种语言环境下验证,该方法为异构医疗系统中临床模型的训练与部署提供了稳健的数据基础。

原文摘要 · Abstract (English)

The widespread adoption of electronic health records has created new opportunities for translational clinical research, yet this promise remains constrained by fragmented data across privacy-siloed institutions and substantial heterogeneity in local coding practices. While privacy-preserving collaborative learning allows institutions to work together without sharing patient-level data, it does not address inconsistencies in how clinical concepts are represented across sites. We introduce a graph-based framework that addresses this gap by treating data harmonization as a scalable representation learning problem. Rather than relying on fixed standards or manual mappings, the framework integrates institution-specific summary statistics from health records, curated biomedical knowledge graphs, and semantic information derived from large language models to learn a shared semantic space. This joint learning approach aligns diverse, site-specific vocabularies while preserving patient privacy. Evaluated across seven institutions and two languages, the framework provides a robust, data-centric foundation for training and deploying clinical models across heterogeneous healthcare systems.

表示学习医疗数据跨机构研究隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。