过参数化下私有学习可免费实现,无需牺牲性能。
Privacy for Free in the Overparameterized Regime
- 在随机特征模型中,大参数量时隐私可无代价获得
- 无论隐私参数ε是否趋近于0,风险均趋于0
- 挑战了过参数化损害私有学习性能的固有认知
差分隐私梯度下降(DP-GD)是训练深度学习模型并保障训练数据隐私的常用算法。过去十年间,研究者们广泛关注其相对于标准梯度下降的性能损失,并在不同学习场景下推导出人口风险超额上界。然而,现有界通常随过参数化(即参数量p大于样本数n)而恶化,这正是当前深度学习实践中的普遍现象。因此,理论空白使从业者缺乏明确指导,有人选择减少可训练参数以提升性能,有人则通过扩大模型规模追求更好结果。本文在带有二次损失的随机特征模型中证明:当参数量p足够大时,隐私可无代价实现,即 |R_P| = o(1),不仅适用于常数级隐私参数ε,也适用于ε → 0的强隐私设置。这一发现挑战了过参数化必然损害私有学习性能的常见认知。
原文摘要 · Abstract (English)
Differentially private gradient descent (DP-GD) is a popular algorithm to train deep learning models with provable guarantees on the privacy of the training data. In the last decade, the problem of understanding its performance cost with respect to standard GD has received remarkable attention from the research community, which formally derived upper bounds on the excess population risk $R_{P}$ in different learning settings. However, existing bounds typically degrade with over-parameterization, i.e., as the number of parameters $p$ gets larger than the number of training samples $n$ -- a regime which is ubiquitous in current deep-learning practice. As a result, the lack of theoretical insights leaves practitioners without clear guidance, leading some to reduce the effective number of trainable parameters to improve performance, while others use larger models to achieve better results through scale. In this work, we show that in the popular random features model with quadratic loss, for any sufficiently large $p$, privacy can be obtained for free, i.e., $\left|R_{P} \right| = o(1)$, not only when the privacy parameter $\varepsilon$ has constant order, but also in the strongly private setting $\varepsilon = o(1)$. This challenges the common wisdom that over-parameterization inherently hinders performance in private learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。