arXiv:2605.02372cs.CRcs.AI2026-05中稿 · the 2nd Internatio…

为联邦学习设计个性化隐私预算,提升医疗数据建模效果

Privacy Preserving Machine Learning Workflow: from Anonymization to Personalized Differential Privacy Budgets in Federated Learning

论文配图:Privacy Preserving Machine Learning Workflow: from Anonymization to Personalized Differential Privacy Budgets in Federated Learning
图 1 · 摘自论文原文
  • 基于重识别风险为不同客户端分配个性化隐私预算
  • 相比固定预算,误差指标降低12.3%与8.7%
  • 可检测客户端漂移,防御数据投毒攻击,适合医疗场景

随着人工智能发展和隐私法规的推动,隐私保护机器学习架构(如联邦学习)日益重要。尽管联邦学习可在不共享数据的前提下训练模型,但仍面临数据完整性和隐私威胁。本文针对敏感表格数据,提出一套包含匿名化与差分隐私技术的全流程隐私保护联邦学习方案。首次正式定义并检测客户端漂移,以缓解投毒攻击。提出基于重识别风险的个性化全局差分隐私预算分配方法。在公开医疗记录数据集上验证,相较于固定隐私预算的全局差分隐私方案,该方法在两项误差指标上分别提升12.3%与8.7%,显著改善模型性能。

原文摘要 · Abstract (English)

The growing development of artificial intelligence based solutions, together with privacy legislation, has driven the rise of the so-called privacy preserving machine learning architectures, such as federated learning. While federated learning enables model training on decentralized data preventing their sharing and centralization, it still faces several challenges related to data integrity and privacy. This paper presents a comprehensive privacy preserving federated learning workflow for sensitive tabular data, including anonymization and differential privacy techniques. We also introduce a formal definition for the concept of client drift, together with ways of detecting it to mitigate poisoning attacks. Then, we detail a complete methodology for assigning personalized privacy budgets for global differential privacy to the different clients participating in the network, based on a re-identification risk metric. The proposed methodology is presented and tested on an openly available dataset of medical records. Within the experimental setup we show that the approach based on personalized budgets, compared to the architecture including global differential privacy with fixed privacy budget, achieves a better model performance in terms of two error metrics.

联邦学习差分隐私医疗数据个性化预算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。