用联邦学习生成胸片报告,保护隐私还能保持高精度。
Privacy-Preserving Chest X-ray Report Generation via Multimodal Federated Learning with ViT and GPT-2
- 用ViT+GPT-2的联邦学习框架,不传原始数据。
- Krum聚合使报告在ROUGE、BLEU等指标上表现最优。
- 适合医疗数据隐私要求高的联合模型开发场景。
从胸片图像自动生成放射科报告在提升诊断效率方面潜力巨大,但传统集中式方法需传输敏感数据,存在隐私风险。为此,本研究提出一种基于Vision Transformer(ViT)和GPT-2的多模态联邦学习框架,使用IU-Xray数据集进行实验。系统采用三种联邦学习聚合策略:FedAvg、Krum Aggregation及新提出的损失感知联邦平均(L-FedAvg)。结果表明,Krum Aggregation在词汇与语义评价指标(如ROUGE、BLEU、BERTScore、RaTEScore)上表现最佳。联邦学习方法生成的报告在临床相关性和语义丰富性上可媲美甚至超越集中式模型。该轻量级且保护隐私的框架为医疗AI协作开发提供了可行路径。
原文摘要 · Abstract (English)
The automated generation of radiology reports from chest X-ray images holds significant promise in enhancing diagnostic workflows while preserving patient privacy. Traditional centralized approaches often require sensitive data transfer, posing privacy concerns. To address this, the study proposes a Multimodal Federated Learning framework for chest X-ray report generation using the IU-Xray dataset. The system utilizes a Vision Transformer (ViT) as the encoder and GPT-2 as the report generator, enabling decentralized training without sharing raw data. Three Federated Learning (FL) aggregation strategies: FedAvg, Krum Aggregation and a novel Loss-aware Federated Averaging (L-FedAvg) were evaluated. Among these, Krum Aggregation demonstrated superior performance across lexical and semantic evaluation metrics such as ROUGE, BLEU, BERTScore and RaTEScore. The results show that FL can match or surpass centralized models in generating clinically relevant and semantically rich radiology reports. This lightweight and privacy-preserving framework paves the way for collaborative medical AI development without compromising data confidentiality.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。