差分隐私在长期数据参与中能提升模型效用,当泄露影响用户留存时。
Performative Privacy: When Differential Privacy Maximizes Utility

- 用差分隐私控制数据泄露,平衡估计误差与用户留存
- 强反馈下,有限隐私预算比完全不私密的估计更优
- 适合关注隐私与数据可持续性的系统设计者
隐私保护学习常基于保护用户数据可维持信任并促进长期参与,从而提升效用。然而这一假设尚未被形式化。与此同时,行为性学习框架研究了部署会改变未来观测数据的学习系统。本文将两者结合,提出行为性隐私(performative privacy):数据泄露会降低未来的参与度。我们研究一个简单模型,其中参与者反复贡献数据用于均值估计,但若数据泄露则可能退出。通过差分隐私机制实现隐私保护,在估计噪声与未来参与之间形成权衡。理论分析和数值实验表明,当泄露与参与之间的反馈回路足够强时,有限隐私预算在长期表现上优于非隐私估计。这首次提供了差分隐私不仅作为保护机制,且从长期效用角度最优的证据。
原文摘要 · Abstract (English)
Privacy-preserving learning is often motivated by the idea that protecting users' data can preserve trust and thus participation, improving utility in the long term. However, this claim has not been formalized so far. In parallel, performative learning provides a framework for studying learning systems whose deployment affects the data they later observe. In this work, we bring these two perspectives together and introduce performative privacy, where data leakage reduces future participation. We study a simple model where agents repeatedly contribute data for mean estimation but may leave the system when their data is leaked. Privacy is implemented through differentially private mechanisms, creating a trade-off between estimation noise and future participation. We show, through a theoretical study of the dynamics and numerical experiments, that a finite privacy budget can outperform non-private estimation in the long term when the feedback loop between leakage and participation is sufficiently strong. This provides first evidence that differential privacy can be optimal not only as a protection mechanism, but also from the perspective of long-term utility.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。