提出噪声校准的推断方法,实现差分隐私下可信赖的统计推断。
Noise-Calibrated Inference from Differentially Private Sufficient Statistics in Exponential Families
- 仅发布经差分隐私处理的充分统计量,再进行噪声校准推断。
- 插件式隐私最大似然估计具渐近正态性与有效置信区间。
- 支持自助法区间计算,适合需严格不确定性量化场景。
许多差分隐私数据发布系统要么输出隐私合成数据,导致推断严重失准;要么仅输出隐私点估计,缺乏可靠的不确定性量化方法。本文针对指数族提出一种清晰且可计算的折中方案:仅发布经过差分隐私处理的充分统计量,随后进行噪声校准的似然推断,并可选地生成参数化合成数据作为后处理。主要贡献包括:(1) 在高斯机制下对截断充分统计量实现近似差分隐私释放的通用方法;(2) 插件式隐私最大似然估计的渐近正态性、显式方差膨胀及有效的沃尔德风格置信区间;(3) 首个一阶等价于插件法但支持自助法区间的噪声感知似然修正;(4) 构建匹配的极小极大下界,证明隐私畸变率不可避免。该理论提供具体设计准则与实用流程,在三个指数族及真实人口普查数据上验证了其在差分隐私合成数据发布中实现可信不确定性量化的有效性。
原文摘要 · Abstract (English)
Many differentially private (DP) data release systems either output DP synthetic data and leave analysts to perform inference as usual, which can lead to severe miscalibration, or output a DP point estimate without a principled way to do uncertainty quantification. This paper develops a clean and tractable middle ground for exponential families: release only DP sufficient statistics, then perform noise-calibrated likelihood-based inference and optional parametric synthetic data generation as post-processing. Our contributions are: (1) a general recipe for approximate-DP release of clipped sufficient statistics under the Gaussian mechanism; (2) asymptotic normality, explicit variance inflation, and valid Wald-style confidence intervals for the plug-in DP MLE; (3) a noise-aware likelihood correction that is first-order equivalent to the plug-in but supports bootstrap-based intervals; and (4) a matching minimax lower bound showing the privacy distortion rate is unavoidable. The resulting theory yields concrete design rules and a practical pipeline for releasing DP synthetic data with principled uncertainty quantification, validated on three exponential families and real census data.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。