用大模型从梯度中还原时空数据,揭示联邦学习隐私风险
Extracting Spatiotemporal Data from Gradients with Large Language Models
- 针对时空数据设计新型梯度反演攻击,可还原用户原始位置
- 引入语言模型辅助搜索,显著提升无先验条件下的重建精度
- 提出自适应防御策略,兼顾隐私保护与模型性能
近期研究发现,敏感用户数据可从梯度更新中重构,破坏联邦学习的隐私承诺。尽管图像数据上的攻击已取得成功,但这些方法难以直接迁移至时空数据领域。为评估时空联邦学习的隐私风险,我们提出针对时空数据的梯度反演攻击(ST-GIA),可成功从梯度中重建原始位置信息。此外,由于缺乏先验知识,现有方法在时空数据上重建效果受限。为此,我们提出ST-GIA+,利用辅助语言模型引导潜在位置搜索,有效恢复原始数据。同时,设计自适应防御策略,动态调整扰动强度,实现不同训练轮次下隐私与效用的更好平衡。在三个真实世界数据集上的大量实验表明,该防御策略能有效保护安全,同时保持时空联邦学习的良好性能。
原文摘要 · Abstract (English)
Recent works show that sensitive user data can be reconstructed from gradient updates, breaking the key privacy promise of federated learning. While success was demonstrated primarily on image data, these methods do not directly transfer to other domains, such as spatiotemporal data. To understand privacy risks in spatiotemporal federated learning, we first propose Spatiotemporal Gradient Inversion Attack (ST-GIA), a gradient attack algorithm tailored to spatiotemporal data that successfully reconstructs the original location from gradients. Furthermore, the absence of priors in attacks on spatiotemporal data has hindered the accurate reconstruction of real client data. To address this limitation, we propose ST-GIA+, which utilizes an auxiliary language model to guide the search for potential locations, thereby successfully reconstructing the original data from gradients. In addition, we design an adaptive defense strategy to mitigate gradient inversion attacks in spatiotemporal federated learning. By dynamically adjusting the perturbation levels, we can offer tailored protection for varying rounds of training data, thereby achieving a better trade-off between privacy and utility than current state-of-the-art methods. Through intensive experimental analysis on three real-world datasets, we reveal that the proposed defense strategy can well preserve the utility of spatiotemporal federated learning with effective security protection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。