提出可量化联邦学习数据重构风险的新指标,助力系统性防御。
From Risk to Resilience: Towards Assessing and Mitigating the Risk of Data Reconstruction Attacks in Federated Learning
- 引入可计算的逆向损失(InvLoss)量化单个数据的重构风险上限。
- 发现攻击风险由模型更新的雅可比矩阵谱特性决定,解释现有防御机制原理。
- 设计自适应噪声防御方法,在不损失精度前提下提升隐私保护能力。
数据重构攻击(DRA)会从联邦学习(FL)中本地客户端的模型更新中推断敏感训练数据,构成严重威胁。尽管已有大量研究,但缺乏理论支撑的风险量化框架,使如何评估和刻画此类攻击风险仍悬而未决。本文提出逆向损失(InvLoss),用于量化给定数据实例与模型下DRA的最大可能有效性,并推导出其紧致且可计算的上界。从三个角度展开分析:第一,揭示攻击风险受交换模型更新或特征嵌入的雅可比矩阵谱特性支配,统一解释多种防御手段的有效性;第二,构建基于InvLoss的攻防无关风险评估器InvRE,实现跨数据实例与模型架构的全面风险评估;第三,提出两种自适应噪声扰动防御策略,在不损害分类准确率的前提下增强隐私保护。在真实世界数据集上的大量实验验证了该框架在系统性风险评估与缓解方面的潜力。
原文摘要 · Abstract (English)
Data Reconstruction Attacks (DRA) pose a significant threat to Federated Learning (FL) systems by enabling adversaries to infer sensitive training data from local clients. Despite extensive research, the question of how to characterize and assess the risk of DRAs in FL systems remains unresolved due to the lack of a theoretically-grounded risk quantification framework. In this work, we address this gap by introducing Invertibility Loss (InvLoss) to quantify the maximum achievable effectiveness of DRAs for a given data instance and FL model. We derive a tight and computable upper bound for InvLoss and explore its implications from three perspectives. First, we show that DRA risk is governed by the spectral properties of the Jacobian matrix of exchanged model updates or feature embeddings, providing a unified explanation for the effectiveness of defense methods. Second, we develop InvRE, an InvLoss-based DRA risk estimator that offers attack method-agnostic, comprehensive risk evaluation across data instances and model architectures. Third, we propose two adaptive noise perturbation defenses that enhance FL privacy without harming classification accuracy. Extensive experiments on real-world datasets validate our framework, demonstrating its potential for systematic DRA risk evaluation and mitigation in FL systems.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。