arXiv:2501.14095stat.MEcs.LG2025-01被引 1

提出更鲁棒的私有均值估计算法,提升数据隐私保护下的统计精度。

Improved subsample-and-aggregate via the private modified winsorized mean

  • 用截尾均值改进子采样-聚合框架中的隐私均值估计
  • 在多种分布下达到最优误差界,对异常值不敏感
  • 适合高维、含噪声或对抗性污染的数据分析场景

我们提出一种一元、差分隐私的均值估计器——私有修改截尾均值,专用于子采样-聚合框架中的聚合器。通过真实数据分析,我们发现常见的一元差分隐私多变量均值估计器即使在大数据集上作为聚合器也表现不佳,由此驱动了本研究。我们证明,修改截尾均值在多个大规模分布类中具有极小最大误差(minimax optimal),且对对抗性污染具有鲁棒性。实验表明,该私有估计器在性能上优于其他私有均值估计方法。我们将修改截尾均值应用于子采样-聚合框架,并推导出使用新聚合器时子采样-聚合估计的有限样本偏差界。该结果揭示两个关键洞察:(i) 最优子样本数量取决于子样本上估计器的偏差;(ii) 子采样-聚合估计器的收敛速率取决于子样本估计器的鲁棒性。

原文摘要 · Abstract (English)

We develop a univariate, differentially private mean estimator, called the private modified winsorized mean, designed to be used as the aggregator in subsample-and-aggregate. We demonstrate, via real data analysis, that common differentially private multivariate mean estimators may not perform well as the aggregator, even in large datasets, motivating our developments.We show that the modified winsorized mean is minimax optimal for several, large classes of distributions, even under adversarial contamination. We also demonstrate that, empirically, the private modified winsorized mean performs well compared to other private mean estimates. We consider the modified winsorized mean as the aggregator in subsample-and-aggregate, deriving a finite sample deviations bound for a subsample-and-aggregate estimate generated with the new aggregator. This result yields two important insights: (i) the optimal choice of subsamples depends on the bias of the estimator computed on the subsamples, and (ii) the rate of convergence of the subsample-and-aggregate estimator depends on the robustness of the estimator computed on the subsamples.

差分隐私均值估计鲁棒统计子采样-聚合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。