提出高效低噪的用户级隐私均值估计方法,适合持续数据流场景。
Matrix Factorization for Practical Continual Mean Estimation Under User-Level Differential Privacy
- 基于矩阵分解设计专用隐私机制,提升估算效率
- 理论证明误差随时间渐进降低,优于已有方法
- 适合需要长期保护用户隐私的数据分析应用
我们研究连续均值估计问题,即数据向量按顺序到达,目标是持续维护运行均值的准确估计。在用户级差分隐私保护下,即使用户贡献多个数据点,也需保护其完整数据集。以往工作主要关注纯差分隐私,但该方法导致噪声过大,限制实际应用。本文改用近似差分隐私,并引入近期矩阵分解机制的进展。我们提出一种专用于均值估计的新矩阵分解方案,兼具高效性与准确性,在用户级差分隐私下的连续均值估计中,实现了渐近更低的均方误差界。
原文摘要 · Abstract (English)
We study continual mean estimation, where data vectors arrive sequentially and the goal is to maintain accurate estimates of the running mean. We address this problem under user-level differential privacy, which protects each user's entire dataset even when they contribute multiple data points. Previous work on this problem has focused on pure differential privacy. While important, this approach limits applicability, as it leads to overly noisy estimates. In contrast, we analyze the problem under approximate differential privacy, adopting recent advances in the Matrix Factorization mechanism. We introduce a novel mean estimation specific factorization, which is both efficient and accurate, achieving asymptotically lower mean-squared error bounds in continual mean estimation under user-level differential privacy.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。