考虑用户数据量不均的隐私均值估计,提出新算法并证明其最优性。
Distribution-Aware Mean Estimation under User-level Local Differential Privacy
- 基于用户数据量分布先验设计均值估计方法
- 理论证明误差上界与下界渐近匹配
- 适合数据量差异大且需用户级隐私保护的场景
我们研究用户级局部差分隐私下的均值估计问题,其中 n 个用户各自拥有来自生成分布 μ 的 m_u 个数据样本,且 m_u 未知但服从已知分布 M。以往工作假设所有用户数据量相同,本文考虑更现实的异构情形。基于分布感知的均值估计算法,我们建立了关于 μ 的最坏情况风险的 M-相关上界,并推导出下界。两者在对数因子内渐近匹配,当所有用户数据量相等时退化为已有结果。
原文摘要 · Abstract (English)
We consider the problem of mean estimation under user-level local differential privacy, where $n$ users are contributing through their local pool of data samples. Previous work assume that the number of data samples is the same across users. In contrast, we consider a more general and realistic scenario where each user $u \in [n]$ owns $m_u$ data samples drawn from some generative distribution $μ$; $m_u$ being unknown to the statistician but drawn from a known distribution $M$ over $\mathbb{N}^\star$. Based on a distribution-aware mean estimation algorithm, we establish an $M$-dependent upper bounds on the worst-case risk over $μ$ for the task of mean estimation. We then derive a lower bound. The two bounds are asymptotically matching up to logarithmic factors and reduce to known bounds when $m_u = m$ for any user $u$.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。