arXiv:2506.17562cs.CVcs.CL2025-06中稿 · IEEE TMI被引 18

用联邦学习让多医院联合训练医疗报告生成大模型,既保护隐私又省通信开销。

LLM-driven Medical Report Generation via Communication-efficient Heterogeneous Federated Learning

  • 用低秩分解压缩参数更新,大幅降低跨中心通信成本。
  • 解决图像特征与报告风格双异构问题,生成报告更准确且符合各院习惯。
  • 适合需要跨机构协作但数据不能外流的医疗AI研发团队使用。

大语言模型在医学报告生成(MRG)中展现出巨大潜力,但其训练需大量医学影像-报告配对数据,这些数据通常分散于多个医疗机构。由于隐私法规限制,集中化数据极为困难,阻碍了模型发展与应用。为此,我们提出FedMRG,首个基于联邦学习(FL)的隐私保护多中心MRG框架,专门应对多模态数据异构下的通信效率难题。首先,通过低秩分解高效分解参数更新,显著降低梯度传输开销,使大模型在带宽受限环境下可行。其次,观察到联邦场景下存在双重异构:不同中心的图像特征差异,以及报告风格和术语偏好各异。为此,我们进一步引入(1)客户端感知对比学习与诊断驱动提示,兼顾全局通用性与局部特异性,保持诊断准确性;(2)解码器中的双适配器互增强机制,协调通用与专用适配器,应对报告风格与术语差异。在构建的FL-MRG基准上广泛评估表明,FedMRG具备良好泛化性与适应性,有望在保障通信效率的同时,利用多中心数据生成临床准确的报告。

原文摘要 · Abstract (English)

LLMs have demonstrated significant potential in Medical Report Generation (MRG), yet their development requires large amounts of medical image-report pairs, which are commonly scattered across multiple centers. Centralizing these data is exceptionally challenging due to privacy regulations, thereby impeding model development and broader adoption of LLM-driven MRG models. To address this challenge, we present FedMRG, the first framework that leverages Federated Learning (FL) to enable privacy-preserving, multi-center development of LLM-driven MRG models, specifically designed to overcome the critical challenge of communication-efficient LLM training under multi-modal data heterogeneity. To start with, our framework tackles the fundamental challenge of communication overhead in FL-LLM tuning by employing low-rank factorization to efficiently decompose parameter updates, significantly reducing gradient transmission costs and making LLM-driven MRG feasible in bandwidth-constrained FL settings. Furthermore, we observed the dual heterogeneity in MRG under the FL scenario: varying image characteristics across medical centers, as well as diverse reporting styles and terminology preferences. To address this, we further enhance FedMRG with (1) client-aware contrastive learning in the MRG encoder, coupled with diagnosis-driven prompts, which capture both globally generalizable and locally distinctive features while maintaining diagnostic accuracy; and (2) a dual-adapter mutual boosting mechanism in the MRG decoder that harmonizes generic and specialized adapters to address variations in reporting styles and terminology. Through extensive evaluation of our established FL-MRG benchmark, we demonstrate the generalizability and adaptability of FedMRG, underscoring its potential in harnessing multi-center data and generating clinically accurate reports while maintaining communication efficiency.

联邦学习医疗AI大模型报告生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。