提出一种在过参数化下保持差分隐私的随机特征模型,兼顾性能与公平性。
Differentially Private Random Feature Model
- 通过输出扰动在过参数化随机特征中实现差分隐私
- 理论证明模型具备隐私保护与泛化误差界,实测性能优于现有方法
- 首次揭示随机特征可缓解隐私机制带来的不公平现象,适合关注公平性的研究者
近年来,设计隐私保护机器学习算法受到广泛关注,尤其在数据包含敏感信息时。差分隐私(DP)是提供隐私保障的数据分析常用机制。本文提出一种差分隐私随机特征模型。随机特征最初用于近似大规模核机,也被用于研究隐私保护核机。我们考虑过参数化情形(特征数大于样本数),其中非私有随机特征模型通过求解最小范数插值问题进行学习,随后应用输出扰动技术生成私有模型。我们证明该方法保持隐私,并推导出其泛化误差上界。据我们所知,这是首个在过参数化设置下研究隐私保护随机特征模型并提供理论保证的工作。我们在合成数据和基准数据集上与文献中其他隐私学习方法进行了实验比较,结果表明本方法在泛化性能上更优。此外,近期研究发现DP机制可能产生并加剧差异影响,即不同群体间预测结果差异显著。我们从理论上和实证上证明,随机特征具有降低差异影响的潜力,从而实现更好的公平性。
原文摘要 · Abstract (English)
Designing privacy-preserving machine learning algorithms has received great attention in recent years, especially in the setting when the data contains sensitive information. Differential privacy (DP) is a widely used mechanism for data analysis with privacy guarantees. In this paper, we produce a differentially private random feature model. Random features, which were proposed to approximate large-scale kernel machines, have been used to study privacy-preserving kernel machines as well. We consider the over-parametrized regime (more features than samples) where the non-private random feature model is learned via solving the min-norm interpolation problem, and then we apply output perturbation techniques to produce a private model. We show that our method preserves privacy and derive a generalization error bound for the method. To the best of our knowledge, we are the first to consider privacy-preserving random feature models in the over-parametrized regime and provide theoretical guarantees. We empirically compare our method with other privacy-preserving learning methods in the literature as well. Our results show that our approach is superior to the other methods in terms of generalization performance on synthetic data and benchmark data sets. Additionally, it was recently observed that DP mechanisms may exhibit and exacerbate disparate impact, which means that the outcomes of DP learning algorithms vary significantly among different groups. We show that both theoretically and empirically, random features have the potential to reduce disparate impact, and hence achieve better fairness.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。