arXiv:2409.09558math.STcs.CR2024-09被引 14

从假设检验视角重新理解差分隐私,统一分析框架更优。

A Statistical Viewpoint on Differential Privacy: Hypothesis Testing, Representation and Blackwell's Theorem

  • 用假设检验重构差分隐私理论基础,揭示其统计本质。
  • 提出f-差分隐私,统一现有定义并支持更精确的隐私边界分析。
  • 适合研究隐私机制设计或机器学习隐私保障的学者参考。

差分隐私因其稳健可靠的保障,被广泛视为隐私保护数据分析的正式标准,已在公共事务、学术界和工业界广泛应用。尽管起源于密码学背景,本文认为差分隐私本质上是纯统计概念。通过利用大卫·布罗克威尔的信息性定理,我们基于已有工作表明,所有差分隐私定义均可从假设检验角度严格推导,证明假设检验不仅是便利工具,更是理解差分隐私的正确语言。这一洞见催生了f-差分隐私的定义,它通过表示定理扩展了其他差分隐私形式。本文综述了将f-差分隐私作为统一框架分析数据与机器学习中隐私边界的技术。应用包括私有深度学习、私有凸优化、洗牌机制及美国人口普查数据,展示了该框架相比传统方法在隐私边界分析上的优势。

原文摘要 · Abstract (English)

Differential privacy is widely considered the formal privacy for privacy-preserving data analysis due to its robust and rigorous guarantees, with increasingly broad adoption in public services, academia, and industry. Despite originating in the cryptographic context, in this review paper we argue that, fundamentally, differential privacy can be considered a \textit{pure} statistical concept. By leveraging David Blackwell's informativeness theorem, our focus is to demonstrate based on prior work that all definitions of differential privacy can be formally motivated from a hypothesis testing perspective, thereby showing that hypothesis testing is not merely convenient but also the right language for reasoning about differential privacy. This insight leads to the definition of $f$-differential privacy, which extends other differential privacy definitions through a representation theorem. We review techniques that render $f$-differential privacy a unified framework for analyzing privacy bounds in data analysis and machine learning. Applications of this differential privacy definition to private deep learning, private convex optimization, shuffled mechanisms, and U.S.\ Census data are discussed to highlight the benefits of analyzing privacy bounds under this framework compared to existing alternatives.

差分隐私统计学假设检验隐私计算

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。