arXiv:2505.01197stat.MLcs.LG2025-05被引 1

提出一种更高效且更准确的私有化自助法,适合大规模数据场景。

Gaussian Differential Private Bootstrap by Subsampling

  • 基于子采样实现非参数自助法,降低重复访问数据需求。
  • 相比传统方法减少噪声添加量,提升统计精度且保持相同隐私水平。
  • 在小样本下表现更优,适用于大规模数据的隐私保护分析。

自助法是量化数据分析不确定性的重要工具。然而,在海量数据上应用自助法时,除了计算成本增加外,差分隐私下的自助法还面临需多次访问数据的问题,导致隐私预算显著增加,进而大幅损失统计精度。为调和隐私与精度的矛盾,过去十年中已有若干基于参数模型的私有自助法被研究,但其有效性依赖于目标量可被识别为模型参数且模型假设近似成立。相比之下,非参数自助法虽在非隐私场景广泛应用,但在隐私条件下研究较少。本文提出一种私有的 $m$ out of $n$ 自助法,并在高斯差分隐私下验证了其一致性与隐私保障。相较于传统的 $n$ out of $n$ 私有自助法,本方法具有三方面优势:第一,计算成本更低,尤其适用于大规模数据;第二,自助迭代中所需额外噪声更少,提升统计精度,同时渐近保证相同隐私水平;第三,有限样本性能明显优于现有方法。

原文摘要 · Abstract (English)

Bootstrap is a common tool for quantifying uncertainty in data analysis. However, besides additional computational costs in the application of the bootstrap on massive data, a challenging problem in bootstrap based inference under Differential Privacy consists in the fact that it requires repeated access to the data. As a consequence, bootstrap based differentially private inference requires a significant increase of the privacy budget, which on the other hand comes with a substantial loss in statistical accuracy. A potential solution to reconcile the conflicting goals of statistical accuracy and privacy is to analyze the data under parametric model assumptions and in the last decade, several parametric bootstrap methods for inference under privacy have been investigated. However, uncertainty quantification by parametric bootstrap is only valid if the the quantities of interest can be identified as the parameters of a statistical model and the imposed model assumptions are (at least approximately) satisfied. An alternative to parametric methods is the empirical bootstrap that is a widely used tool for non-parametric inference and well studied in the non-private regime. However, under privacy, less insight is available. In this paper, we propose a private empirical $m$ out of $n$ bootstrap and validate its consistency and privacy guarantees under Gaussian Differential Privacy. Compared to the the private $n$ out of $n$ bootstrap, our approach has several advantages. First, it comes with less computational costs, in particular for massive data. Second, the proposed procedure needs less additional noise in the bootstrap iterations, which leads to an improved statistical accuracy while asymptotically guaranteeing the same level of privacy. Third, we demonstrate much better finite sample properties compared to the currently available procedures.

差分隐私自助法统计推断子采样

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。