公共数据助训的联邦蒸馏仍会泄露客户隐私,需警惕。
Unveiling Client Privacy Leakage from Public Dataset Usage in Federated Distillation
- 利用客户端在公开数据上的推理结果,反推私有数据分布
- 提出新型攻击方法,在低误报率下实现高查准率的成员推断
- 首次系统揭示公共数据辅助联邦蒸馏的隐私风险,适合关注隐私安全的研究者
联邦蒸馏(FD)作为一种流行的联邦学习框架,使客户端能在不共享私有数据的情况下协作训练模型。公共数据辅助联邦蒸馏(PDA-FD)通过利用公开数据集进行知识传递,已广泛采用。尽管相比传统联邦学习更具隐私性,本文首次在诚实但好奇服务器假设下,全面分析了PDA-FD的隐私风险。我们证明,服务器可利用客户端在公开数据集上的推理结果,提取两类关键隐私信息:私有训练数据的标签分布和成员身份信息。为此,提出两种针对PDA-FD场景设计的新攻击方法:标签分布推断攻击与基于似然比攻击(LiRA)的创新成员推断方法。在三种代表性框架(FedMD、DS-FL、Cronus)上评估显示,标签分布攻击达到最小KL散度,成员推断攻击在低假阳性率下保持高真阳性率,性能达当前最优。研究揭示现有PDA-FD框架存在显著隐私漏洞,亟需更鲁棒的隐私保护机制。
原文摘要 · Abstract (English)
Federated Distillation (FD) has emerged as a popular federated training framework, enabling clients to collaboratively train models without sharing private data. Public Dataset-Assisted Federated Distillation (PDA-FD), which leverages public datasets for knowledge sharing, has become widely adopted. Although PDA-FD enhances privacy compared to traditional Federated Learning, we demonstrate that the use of public datasets still poses significant privacy risks to clients' private training data. This paper presents the first comprehensive privacy analysis of PDA-FD in presence of an honest-but-curious server. We show that the server can exploit clients' inference results on public datasets to extract two critical types of private information: label distributions and membership information of the private training dataset. To quantify these vulnerabilities, we introduce two novel attacks specifically designed for the PDA-FD setting: a label distribution inference attack and innovative membership inference methods based on Likelihood Ratio Attack (LiRA). Through extensive evaluation of three representative PDA-FD frameworks (FedMD, DS-FL, and Cronus), our attacks achieve state-of-the-art performance, with label distribution attacks reaching minimal KL-divergence and membership inference attacks maintaining high True Positive Rates under low False Positive Rate constraints. Our findings reveal significant privacy risks in current PDA-FD frameworks and emphasize the need for more robust privacy protection mechanisms in collaborative learning systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。