通过联邦学习联合多中心数据,提升罕见病诊断准确率。
Training Together, Diagnosing Better: Federated Learning for Collagen VI-Related Dystrophies
- 在不共享原始数据前提下,跨机构协作训练诊断模型
- 对三种致病机制分类达F1分数0.82,显著优于单机构模型
- 适合罕见病研究、医疗数据隐私敏感场景使用
将机器学习应用于罕见病如胶原蛋白VI相关营养不良(COL6-RD)的诊断,受限于数据稀缺与分散。跨医院、机构或国家的数据扩展面临严重隐私、监管和物流障碍。联邦学习(FL)提供了一种可行方案,可在保持患者数据本地化的同时实现协同模型训练。本文报告了一项基于Sherpa.ai FL平台的全球性联邦学习倡议,利用两个国际组织的分布式数据集,对患者来源成纤维细胞培养物的胶原蛋白VI免疫荧光显微图像进行COL6-RD诊断。该方法成功构建出可将患者图像分类为三类主要致病机制(外显子跳跃、甘氨酸替代、假外显子插入)的ML模型,取得F1分数0.82,显著高于单机构模型(0.57–0.75)。结果表明,相比孤立机构模型,联邦学习显著提升了诊断效能与泛化能力。本方法不仅有助于更精准诊断,还可支持意义未明变异的解读,并指导测序策略以发现新致病变异。
原文摘要 · Abstract (English)
The application of Machine Learning (ML) to the diagnosis of rare diseases, such as collagen VI-related dystrophies (COL6-RD), is fundamentally limited by the scarcity and fragmentation of available data. Attempts to expand sampling across hospitals, institutions, or countries with differing regulations face severe privacy, regulatory, and logistical obstacles that are often difficult to overcome. The Federated Learning (FL) provides a promising solution by enabling collaborative model training across decentralized datasets while keeping patient data local and private. Here, we report a novel global FL initiative using the Sherpa.ai FL platform, which leverages FL across distributed datasets in two international organizations for the diagnosis of COL6-RD, using collagen VI immunofluorescence microscopy images from patient-derived fibroblast cultures. Our solution resulted in an ML model capable of classifying collagen VI patient images into the three primary pathogenic mechanism groups associated with COL6-RD: exon skipping, glycine substitution, and pseudoexon insertion. This new approach achieved an F1-score of 0.82, outperforming single-organization models (0.57-0.75). These results demonstrate that FL substantially improves diagnostic utility and generalizability compared to isolated institutional models. Beyond enabling more accurate diagnosis, we anticipate that this approach will support the interpretation of variants of uncertain significance and guide the prioritization of sequencing strategies to identify novel pathogenic variants.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。