用外部数据提升肺癌患者二次癌症预测准确率
LF2L: Loss Fusion Horizontal Federated Learning Across Heterogeneous Feature Spaces Using External Datasets Effectively: A Case Study in Second Primary Cancer Prediction
- 通过损失融合联邦学习,跨机构协作不共享数据
- 在台湾和美国数据上实现AUROC与AUPRC显著提升
- 适合医疗数据隐私敏感场景下的模型优化
二次原发性癌症(SPC)是既往癌症患者新发的另一种癌症,随着癌症存活率提高日益受到关注。早期预测对及时临床干预至关重要。本研究聚焦台湾医院治疗的肺癌幸存者,受限于本地数据规模小、地理范围有限,传统机器学习效果不佳。为此,我们引入美国基于监测、流行病学与结果统计(SEER)项目的公开外部数据,大幅扩充数据多样性和规模。但多源数据存在特征不一致和隐私限制等挑战。本文提出一种损失融合横向联邦学习(LF2L)框架,无需共享数据即可实现跨机构协作。通过同时利用共通与独特特征,并以共享损失机制平衡贡献,显著提升SPC预测性能。实验表明,相较于本地、横向联邦及集中式学习基线,本方法在AUROC和AUPRC上均有统计学显著提升。这凸显了不仅需获取外部数据,更要有效利用以增强真实世界临床模型性能。
原文摘要 · Abstract (English)
Second primary cancer (SPC), a new cancer in patients different from previously diagnosed, is a growing concern due to improved cancer survival rates. Early prediction of SPC is essential to enable timely clinical interventions. This study focuses on lung cancer survivors treated in Taiwanese hospitals, where the limited size and geographic scope of local datasets restrict the effectiveness and generalizability of traditional machine learning approaches. To address this, we incorporate external data from the publicly available US-based Surveillance, Epidemiology, and End Results (SEER) program, significantly increasing data diversity and scale. However, the integration of multi-source datasets presents challenges such as feature inconsistency and privacy constraints. Rather than naively merging data, we proposed a loss fusion horizontal federated learning (LF2L) framework that can enable effective cross-institutional collaboration while preserving institutional privacy by avoiding data sharing. Using both common and unique features and balancing their contributions through a shared loss mechanism, our method demonstrates substantial improvements in the prediction performance of SPC. Experiment results show statistically significant improvements in AUROC and AUPRC when compared to localized, horizontal federated, and centralized learning baselines. This highlights the importance of not only acquiring external data but also leveraging it effectively to enhance model performance in real-world clinical model development.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。