攻击隐私保护的查询系统,用少量恶意数据放大结果偏差5-10倍
Data Poisoning Attacks to Locally Differentially Private Range Query Protocols
- 设计针对树结构与网格结构的最优投毒方法,伪造数据一致且隐蔽
- 实测仅操控少量用户就使任意区间查询估计值放大5-10倍
- 揭示常见后处理机制反而助涨攻击效果,适合研究隐私安全者阅读
本地差分隐私(LDP)被广泛用于去中心化数据收集中的用户隐私保护。然而,近期研究发现LDP协议易受数据投毒攻击,即恶意用户通过篡改上报数据来扭曲聚合结果。本文首次系统研究针对LDP范围查询协议的数据投毒攻击,聚焦树结构与网格结构两种方案。我们识别出三方面挑战:构造一致有效的虚假数据、跨层级或网格保持数据一致性、规避服务器检测。为解决前两项挑战,提出可证明最优的新攻击方法——树基攻击与网格基攻击,能高效操纵范围查询结果。关键发现是:LDP范围查询协议中常见的后处理步骤Norm-Sub会极大增强攻击效果。此外,针对潜在防御措施,我们设计自适应攻击以绕过检测。通过理论分析与在合成及真实数据集上的大量实验验证,结果表明所提攻击能仅通过操纵少量用户,显著放大任意范围查询的估计值,攻击影响力可达正常用户的5-10倍。
原文摘要 · Abstract (English)
Local Differential Privacy (LDP) has been widely adopted to protect user privacy in decentralized data collection. However, recent studies have revealed that LDP protocols are vulnerable to data poisoning attacks, where malicious users manipulate their reported data to distort aggregated results. In this work, we present the first study on data poisoning attacks targeting LDP range query protocols, focusing on both tree-based and grid-based approaches. We identify three key challenges in executing such attacks, including crafting consistent and effective fake data, maintaining data consistency across levels or grids, and preventing server detection. To address the first two challenges, we propose novel attack methods that are provably optimal, including a tree-based attack and a grid-based attack, designed to manipulate range query results with high effectiveness. \textbf{Our key finding is that the common post-processing procedure, Norm-Sub, in LDP range query protocols can help the attacker massively amplify their attack effectiveness.} In addition, we study a potential countermeasure, but also propose an adaptive attack capable of evading this defense to address the third challenge. We evaluate our methods through theoretical analysis and extensive experiments on synthetic and real-world datasets. Our results show that the proposed attacks can significantly amplify estimations for arbitrary range queries by manipulating a small fraction of users, providing 5-10x more influence than a normal user to the estimation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。