arXiv:2505.22108cs.LGcs.AI2025-05被引 1

根据机构合规程度动态分配隐私噪声,让医疗联邦学习更包容且不损失性能。

Inclusive Federated Learning Through Compliance-Weighted Noise Allocation in Healthcare AI

  • 按机构合规度调整隐私噪声,避免对高合规方不公平惩罚。
  • 引入12个低合规机构后模型准确率提升最高达17个百分点。
  • 实现可审计的个性化隐私保护,适合医疗数据协作场景。

联邦学习(FL)可在不集中患者数据的前提下协同训练临床AI模型,但其应用受限于隐私担忧、机构合规差异及资源不均;传统差分隐私(DP)对所有客户端统一加噪,损害了高合规或资源匮乏机构的利益。本文提出一种合规感知的联邦学习框架,将差分隐私适配至机构合规水平,使低合规机构参与时不影响其他方性能。通过与HIPAA、GDPR、NIST、ISO、HL7/FHIR对齐的合规评分工具,将每家机构的评分映射为每轮的高斯噪声尺度,应用于小规模聚合器数据集上的服务器端DP-SGD。形式化的(ε,δ)边界适用于半诚实聚合器下的聚合器数据集;客户端级差分隐私需安全聚合(未来工作)。在PneumoniaMNIST和BreastMNIST数据集上评估五种联邦策略(16客户端,50轮,5次随机种子),聚合器数据集的累计ε值分别为513(肺炎)和1434(乳腺),δ=10⁻⁵。结果显示,相比仅允许高合规机构参与的基线(实验4),包含12个低合规机构(实验1)使乳腺癌模型准确率提升+4.5(FedAvg)、+6.8(FedMedian)、+5.2(FedProx)、+1.6(FedYogi)、-4.1(FedAdam)百分点(合并平均+2.8个百分点,未达显著性,n=5;单配置最高+17个百分点);合规加权分配在平均噪声相当的情况下,性能无损耗(仅+0.1个百分点),首轮噪声代价为1.3个百分点(乳腺)和2.5个百分点(肺炎,FedAvg)。结论:基于合规性的服务器端差分隐私使低合规机构可参与联邦学习,且不牺牲模型性能,提供可审计的逐机构噪声控制;形式化保障作用于聚合器数据集,客户端级隐私需依赖安全聚合。

原文摘要 · Abstract (English)

Background: Federated learning (FL) enables collaborative training of clinical AI models without centralizing patient data, but adoption is limited by privacy concerns, heterogeneous institutional compliance, and resource disparities; standard differential privacy (DP) applies uniform noise to all clients, penalizing well-compliant or under-resourced institutions. Objective: We introduce a compliance-aware FL framework that adapts DP to institutional compliance, letting lower-compliance sites participate without uniformly penalizing others. Methods: A compliance scoring tool aligned with HIPAA, GDPR, NIST, ISO, and HL7/FHIR maps each client score to a per-step Gaussian noise scale for server-side DP-SGD on a small aggregator dataset. The formal $(ε,δ)$ bound applies to the aggregator dataset under a semi-honest aggregator; client-level DP needs secure aggregation (future work). We evaluate five FL strategies on PneumoniaMNIST and BreastMNIST (16 clients, 50 rounds, five seeds); the cumulative aggregator-dataset $ε$ is about 1434 (Breast) and 513 (Pneumonia) at $δ=10^{-5}$. Results: Including 12 lower-compliance clients (Experiment 1) versus a compliant-only baseline (Experiment 4) changed BreastMNIST accuracy by +4.5 (FedAvg), +6.8 (FedMedian), +5.2 (FedProx), +1.6 (FedYogi), and -4.1 (FedAdam) percentage points (pooled +2.8 pp; not significant at n=5; up to +17 pp per configuration); compliance-weighted allocation matched uniform server-side DP at equal mean noise (+0.1 pp), carrying no utility penalty, and first-round noise cost 1.3 pp (Breast) and 2.5 pp (Pneumonia, FedAvg). Conclusions: Compliance-weighted server-side DP lets lower-compliance institutions join FL without degrading performance, giving auditable per-site noise control at no utility cost; formal guarantees apply to the aggregator dataset, with client-level DP requiring secure aggregation.

联邦学习医疗AI差分隐私隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。