提出新攻击方法,精准推断联邦学习中客户端标签分布。
Whispers of Data: Unveiling Label Distributions in Federated Learning Through Virtual Client Simulation
- 用虚拟客户端模拟目标客户端数据特征
- 通过时间泛化差异预测标签比例,准确率超现有方法
- 在差分隐私防御下仍有效,适合安全评估场景
联邦学习允许多个地理分散的客户端协作训练全局模型,无需共享数据。但易受标签推断攻击。现有研究对目标客户端设置敏感,且在防御策略下表现不佳。本文提出一种新型标签分布推断攻击,具备稳定性和适应性。具体而言,先估计目标客户端数据规模,构建多个针对性虚拟客户端;再量化各虚拟客户端的类别时间泛化能力,利用其变化特征训练推断模型,以预测目标客户端的标签分布比例。在MNIST、Fashion-MNIST、FER2013和AG-News等多个数据集上验证,结果表明该方法优于当前最优技术。尤其在差分隐私防御机制下仍保持有效性,凸显其在真实场景中的应用潜力。
原文摘要 · Abstract (English)
Federated Learning enables collaborative training of a global model across multiple geographically dispersed clients without the need for data sharing. However, it is susceptible to inference attacks, particularly label inference attacks. Existing studies on label distribution inference exhibits sensitive to the specific settings of the victim client and typically underperforms under defensive strategies. In this study, we propose a novel label distribution inference attack that is stable and adaptable to various scenarios. Specifically, we estimate the size of the victim client's dataset and construct several virtual clients tailored to the victim client. We then quantify the temporal generalization of each class label for the virtual clients and utilize the variation in temporal generalization to train an inference model that predicts the label distribution proportions of the victim client. We validate our approach on multiple datasets, including MNIST, Fashion-MNIST, FER2013, and AG-News. The results demonstrate the superiority of our method compared to state-of-the-art techniques. Furthermore, our attack remains effective even under differential privacy defense mechanisms, underscoring its potential for real-world applications.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。