研究深度学习生成轨迹时差分隐私带来的性能损失
What is the Cost of Differential Privacy for Deep Learning-Based Trajectory Generation?
- 用差分隐私训练生成模型,牺牲部分精度以保障隐私
- 大数据集下仍保留一定生成质量,小数据集则性能下降明显
- 扩散模型最优但无隐私保障,加了隐私保护后生成对抗网络表现最好
位置轨迹虽具价值,却可能泄露敏感个人信息。差分隐私(DP)提供形式化保护,但实现良好效用-隐私平衡仍具挑战。现有方法多采用深度学习生成模型合成轨迹,但缺乏正式隐私保证,且生成过程依赖真实数据的条件信息。本文在两个数据集上,通过十一个效用指标,针对三个研究问题展开分析:(1)评估标准差分隐私训练方法DP-SGD对先进生成模型效用的影响;(2)针对仅适用于无条件生成的DP-SGD,提出一种新型条件生成差分隐私机制,提供形式化保障并评估其效用影响;(3)分析扩散模型、变分自编码器和生成对抗网络三类模型对效用-隐私权衡的影响。结果表明,DP-SGD显著降低性能,但在数据量足够大时仍可保持部分效用。所提机制提升训练稳定性,尤其与DP-SGD结合使用时,在小型数据集及不稳定的模型(如GAN)上效果更佳。无隐私保障下扩散模型表现最佳,但加入DP-SGD后生成对抗网络性能最优,说明非私有最优模型未必在形式保障下依然最优。结论:当前差分隐私轨迹生成仍具挑战,形式保障仅在大数据集和受限场景中可行。
原文摘要 · Abstract (English)
While location trajectories offer valuable insights, they also reveal sensitive personal information. Differential Privacy (DP) offers formal protection, but achieving a favourable utility-privacy trade-off remains challenging. Recent works explore deep learning-based generative models to produce synthetic trajectories. However, current models lack formal privacy guarantees and rely on conditional information derived from real data during generation. This work investigates the utility cost of enforcing DP in such models, addressing three research questions across two datasets and eleven utility metrics. (1) We evaluate how DP-SGD, the standard DP training method for deep learning, affects the utility of state-of-the-art generative models. (2) Since DP-SGD is limited to unconditional models, we propose a novel DP mechanism for conditional generation that provides formal guarantees and assess its impact on utility. (3) We analyse how model types - Diffusion, VAE, and GAN - affect the utility-privacy trade-off. Our results show that DP-SGD significantly impacts performance, although some utility remains if the datasets is sufficiently large. The proposed DP mechanism improves training stability, particularly when combined with DP-SGD, for unstable models such as GANs and on smaller datasets. Diffusion models yield the best utility without guarantees, but with DP-SGD, GANs perform best, indicating that the best non-private model is not necessarily optimal when targeting formal guarantees. In conclusion, DP trajectory generation remains a challenging task, and formal guarantees are currently only feasible with large datasets and in constrained use cases.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。