提出一种新型隐私保护自助法,通过分组加权提升统计推断的隐私性。
Private Generative Bootstrap via Blocking

- 将个体分组后统一赋权,隐藏个人贡献以增强差分隐私
- 训练时加入校准噪声学习映射,后续采样无需额外隐私开销
- 适用于无需假设数据生成模型的隐私统计推断,适合政策评估等场景
随着人工智能系统越来越多地访问个人数据,报告统计结果时保护隐私至关重要,同时需对结果不确定性进行隐私化处理。为此,本文采用贝叶斯无似然框架,使后验模拟过程具备隐私性。提出一种基于分组策略的新私有化贝叶斯自助法(PGBB):不为每个个体分配独立随机权重,而是将个体随机分组,每组赋予单一权重。通过隐藏组内个体贡献,强化差分隐私保障。利用摊销推理,将私有学习与后验采样解耦。通过在训练中添加校准噪声,学习从观测权重到后验样本的前向映射,后续采样无需额外隐私和计算成本。理论证明了差分隐私保证,分析了收敛到非私有分组自助目标的性能,并量化了普通与分组贝叶斯自助后验之间的差异。此外,推导出无需数据即可调参的块狄利克雷浓度参数设置方法,可渐近恢复后验发散性。还证明单次拟合的PGBB可同时支持一系列基于损失的决策规则,且无需额外隐私代价。在模拟实验及美国人口普查教育回报、出生体重分位数等实际应用中,PGBB展现出竞争力的隐私不确定性量化能力,优于依赖数据生成模型设定的现有私有贝叶斯方法。
原文摘要 · Abstract (English)
With AI systems gaining more access to individuals' information, it is important to protect privacy when reporting statistical answers. Equally important is to privatize the reporting of uncertainty in such answers. To this end, we adopt a Bayesian likelihood-free framework and make simulation from the posterior private. In particular, we propose a new private instantiation of the Bayesian bootstrap using a blocking strategy. Rather than assigning idiosyncratic random weights to each individual, we randomly group individuals and assign a single weight to each group. By concealing individuals' contributions within a group, we fortify differential privacy gates. We harness amortized inference that decouples private learning from posterior sampling. A push-forward map from observation weights to posterior samples is learned privately by adding calibrated noise during training. Subsequent posterior draws require no additional privacy and computation budget. We call the resulting method the Private Generative Bayesian Bootstrap (PGBB). We establish a differential privacy guarantee, analyze convergence to the non-private blocked-bootstrap target, and quantify the discrepancy between the ordinary and blocked Bayesian-bootstrap posteriors. In addition, we derive data-free tuning of the block Dirichlet concentration parameter that restores posterior dispersion asymptotically. We also show a single fit of PGBB can support a family of loss-based decision rules simultaneously without additional privacy cost. In simulations and in applications to U.S. Census returns to schooling and U.S. natality birthweight quantiles, PGBB gives competitive private uncertainty quantification and improves over private Bayesian alternatives that require a specified data-generating model in common settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。