arXiv:2503.14709cs.LGcs.CR2025-03被引 3

利用不可靠的外部数据,让私有分布检测更省样本。

Better Private Distribution Testing by Leveraging Unverified Auxiliary Data

  • 用不靠谱的公开数据辅助私有检测,降低隐私成本。
  • 三种经典测试的样本量随外部数据可信度线性下降。
  • 适合需要隐私保护且有历史数据的科研与工程场景。

我们将增强型分布测试框架(Aliakbarpour 等,NeurIPS 2024)拓展至差分隐私设置。该框架刻画了数据分析师需在敏感数据上执行假设检验,但可利用关于数据分布的先验知识(公共但可能错误或不可信)的场景。我们设计了适用于统一性、身份性和接近性检测这三类典型分布测试任务的私有算法,其样本复杂度随所声称的辅助信息质量平滑下降。我们还给出了信息论下界,证明这些算法的样本复杂度在对数因子内是最优的。

原文摘要 · Abstract (English)

We extend the framework of augmented distribution testing (Aliakbarpour, Indyk, Rubinfeld, and Silwal, NeurIPS 2024) to the differentially private setting. This captures scenarios where a data analyst must perform hypothesis testing tasks on sensitive data, but is able to leverage prior knowledge (public, but possibly erroneous or untrusted) about the data distribution. We design private algorithms in this augmented setting for three flagship distribution testing tasks, uniformity, identity, and closeness testing, whose sample complexity smoothly scales with the claimed quality of the auxiliary information. We complement our algorithms with information-theoretic lower bounds, showing that their sample complexity is optimal (up to logarithmic factors).

隐私计算分布检测差分隐私辅助数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。