arXiv:2505.20020cs.LGcs.SE2025-05被引 14

用知识图谱和大模型统一医疗数据,让多机构协作训练更安全

Ontology- and LLM-based Data Harmonization for Federated Learning in Healthcare

  • 结合医学本体与大模型进行语义对齐
  • 在真实医疗项目中实现跨机构数据统一
  • 适合医疗联邦学习中的隐私保护场景

电子健康记录(EHR)的兴起为医学研究带来了新机遇,但隐私法规和数据异质性仍是大规模机器学习的主要障碍。联邦学习(FL)可在不共享原始数据的情况下实现协作建模,但仍面临多样临床数据难以对齐的问题。本文提出一种两阶段数据对齐策略,融合医学本体与大语言模型(LLMs),支持医疗领域安全、隐私保护的联邦学习,在一个真实项目中实现了EHR数据的语义映射,验证了其有效性。

原文摘要 · Abstract (English)

The rise of electronic health records (EHRs) has unlocked new opportunities for medical research, but privacy regulations and data heterogeneity remain key barriers to large-scale machine learning. Federated learning (FL) enables collaborative modeling without sharing raw data, yet faces challenges in harmonizing diverse clinical datasets. This paper presents a two-step data alignment strategy integrating ontologies and large language models (LLMs) to support secure, privacy-preserving FL in healthcare, demonstrating its effectiveness in a real-world project involving semantic mapping of EHR data.

联邦学习医疗AI数据对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。