揭示隐私与泛化关系,给出DP-SGD的线性最大信息界。
From Privacy to Generalization: Linear Max-Information Bounds for DP-SGD
- 基于线性规模的最大信息界分析DP-SGD的泛化能力。
- 首次获得可显式控制超参数的泛化误差上界。
- 适用于希望理解隐私-泛化权衡的研究者。
理解泛化与隐私之间的关系仍是现代机器学习理论中的核心挑战,尤其针对通过差分隐私随机梯度下降(DP-SGD)训练的深度网络。本文在这一长期开放问题上取得进展,证明了DP-SGD的近似最大信息量存在有限样本上界,其量级与Dwork等人(2015)关于ε-差分隐私算法的经典结果相当,即最多随数据集规模线性增长。由此我们推导出一个通用的PAC-Bayes泛化界,其中所需先验分布可通过DP-SGD学习;同时获得了直接针对DP-SGD训练模型的泛化界,其复杂度项完全显式,并由优化超参数可控。
原文摘要 · Abstract (English)
Understanding the relationship between generalization and privacy remains a central challenge in modern machine learning theory, particularly for deep networks trained by variants of differentially private stochastic gradient descent (DP-SGD). In this work we make progress on this persistent open problem by proving a finite-sample bound on the approximate max-information of DP-SGD that exhibits scaling properties comparable with (Dwork et al, 2015)'s classic result for $ε$-differentially private algorithms, namely at most linear in the dataset size. From our result we obtain a general-purpose PAC-Bayes generalization bound in which the necessary prior distribution can be learned by DP-SGD, as well as a generalization bound for DP-SGD-trained models themselves, with a complexity term that is fully explicit and controlled by the optimization hyperparameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。