提出局部最优私有采样方法,提升隐私保护下的数据生成质量。
Locally Optimal Private Sampling: Beyond the Global Minimax
- 从局部分布邻域出发,构建更贴近实际的隐私采样框架。
- 证明局部极小极大风险由全局风险决定,且存在闭式最优解。
- 适用于有公开数据的场景,实测优于传统全局最优方法。
我们研究在局部差分隐私(LDP)约束下从分布中采样的问题。给定一个私有分布 $P \in \mathcal{P}$,目标是生成一个与 $P$ 在 $f$-散度意义下接近的分布样本,同时满足 LDP 要求。该任务捕捉了在强隐私保障下生成真实数据的根本挑战。以往工作聚焦于分布类上的全局极小极大最优性,本文则采用局部视角:考察固定分布 $P_0$ 附近的极小极大风险,并刻画其精确值,该值依赖于 $P_0$ 和隐私水平。主要结果表明,当分布类 $\mathcal{P}$ 局限于 $P_0$ 的邻域时,局部极小极大风险由全局极小极大风险决定。为此,我们(1)将先前工作从纯 LDP 推广到更一般的函数型 LDP 框架;(2)证明全局最优函数型 LDP 采样器在受限于 $P_0$ 邻域时即为局部最优。进一步,我们推导出不依赖 $f$-散度选择的局部极小极大最优采样器的闭式表达。此外,我们指出该局部框架自然适用于包含公共数据的私有采样场景,其中公共数据分布由 $P_0$ 表示。实验比较表明,所提局部最优采样器在多个设置下持续优于现有全局方法。
原文摘要 · Abstract (English)
We study the problem of sampling from a distribution under local differential privacy (LDP). Given a private distribution $P \in \mathcal{P}$, the goal is to generate a single sample from a distribution that remains close to $P$ in $f$-divergence while satisfying the constraints of LDP. This task captures the fundamental challenge of producing realistic-looking data under strong privacy guarantees. While prior work by Park et al. (NeurIPS'24) focuses on global minimax-optimality across a class of distributions, we take a local perspective. Specifically, we examine the minimax risk in a neighborhood around a fixed distribution $P_0$, and characterize its exact value, which depends on both $P_0$ and the privacy level. Our main result shows that the local minimax risk is determined by the global minimax risk when the distribution class $\mathcal{P}$ is restricted to a neighborhood around $P_0$. To establish this, we (1) extend previous work from pure LDP to the more general functional LDP framework, and (2) prove that the globally optimal functional LDP sampler yields the optimal local sampler when constrained to distributions near $P_0$. Building on this, we also derive a simple closed-form expression for the locally minimax-optimal samplers which does not depend on the choice of $f$-divergence. We further argue that this local framework naturally models private sampling with public data, where the public data distribution is represented by $P_0$. In this setting, we empirically compare our locally optimal sampler to existing global methods, and demonstrate that it consistently outperforms global minimax samplers.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。