arXiv:2503.20107eess.IVphysics.med-ph2025-03被引 6

联邦学习让多家医院联合训练影像模型,不交换数据也能提升诊断准确率。

Federated Learning: A new frontier in the exploration of multi-institutional medical imaging data

  • 各医院本地训练模型,仅上传参数更新,保护患者隐私。
  • 可整合多源异构医学影像数据,提升模型泛化能力。
  • 适合医疗数据隐私要求高、跨机构合作难的场景。

人工智能已深刻改变医学影像领域,推动现代计算机辅助医疗系统的技术革新。然而,深度学习系统普遍需要大量数据以实现有效知识提取与泛化。由于伦理协议签署、数据采集流程建立及妥善管理(尤其是匿名化)耗时费力,获取大规模数据面临挑战。深度学习领域的关键难题之一是整合来自不同硬件厂商、多样采集协议、实验设置乃至操作者差异的数据。本文综述了联邦学习(FL)概念,该方法可在多个医疗机构间协作训练深度学习模型,无需集中数据。相比中心化方案,去中心化的联邦学习在保护各机构数据隐私的同时完成模型训练。我们系统梳理了联邦学习的基本原则,全面回顾通用与专用医学影像聚合与学习算法,支持生成全局泛化模型。深入分析构建联邦学习系统所面临的挑战,包括机构间数据与模型异质性、对数据隐私攻击的鲁棒性,以及计算与通信资源差异导致的系统效率问题。最后,介绍当前主流开源框架,展示实际应用案例,并展望该快速发展的领域的未来方向。

原文摘要 · Abstract (English)

Artificial intelligence has transformed the perspective of medical imaging, leading to a genuine technological revolution in modern computer-assisted healthcare systems. However, ubiquitously featured deep learning (DL) systems require access to a considerable amount of data, facilitating proper knowledge extraction and generalization. Access to such extensive resources may be hindered due to the time and effort required to convey ethical agreements, set up and carry the acquisition procedures through, and manage the datasets adequately with a particular emphasis on proper anonymization. One of the pivotal challenges in the DL field is data integration from various sources acquired using different hardware vendors, diverse acquisition protocols, experimental setups, and even inter-operator variabilities. In this paper, we review the federated learning (FL) concept that fosters the integration of large-scale heterogeneous datasets from multiple institutions in training DL models. In contrast to a centralized approach, the decentralized FL procedure promotes training DL models while preserving data privacy at each institution involved. We formulate the FL principle and comprehensively review general and specialized medical imaging aggregation and learning algorithms, enabling the generation of a globally generalized model. We meticulously go through the challenges in constructing FL-based systems, such as data and model heterogeneities across the institutions, resilience to potential attacks on data privacy, and the variability in computational and communication resources among the entangled sites that might induce efficiency issues of the entire system. Finally, we explore the up-to-date open frameworks for rapid FL-based algorithm prototyping, comprehensively present real-world implementations of FL systems and shed light on future directions in this intensively growing field.

联邦学习医学影像数据隐私多中心研究

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。