提出可高效验证的私有机器学习方法,验证成本低于训练成本。
Efficient Public Verification of Private ML via Regularization
- 通过正则化优化序列实现最优隐私-效用权衡
- 验证所需计算量远低于模型训练开销
- 适合需要可信隐私保障的大规模数据场景
使用差分隐私(DP)训练可确保数据提供者不会被模型识别。然而,数据提供者及公众缺乏高效验证模型是否满足DP保障的方法。当前算法的DP验证计算量与训练计算量成正比。本文设计首个具有近似最优隐私-效用权衡的DP算法,其DP保障可比训练成本更低地进行验证。聚焦于差分隐私凸优化(DP-SCO),我们通过私有最小化一系列正则化目标,仅使用标准的DP组合界,即可获得紧致的隐私-效用权衡。关键在于,该方法的验证计算量远低于训练开销。这使得首个已知的、在大规模数据上验证成本显著低于训练成本的近最优DP-SCO算法成为可能。
原文摘要 · Abstract (English)
Training with differential privacy (DP) guarantees dataset members that they cannot be identified by users of the released model. However, those data providers, and, in general, the public, lack methods to efficiently verify that models trained on their data satisfy DP guarantees. The amount of compute needed to verify DP guarantees for current algorithms scales with the amount of computation required to train the model. In this paper we design the first DP algorithm with near optimal privacy-utility trade-offs but whose DP guarantees can be verified cheaper than training. We focus on DP stochastic convex optimization (DP-SCO), where optimal privacy-utility trade-offs are known. Here we show we can obtain tight privacy-utility trade-offs by privately minimizing a series of regularized objectives and only using the standard DP composition bound. Crucially, this method can be verified with much less compute than training. This leads to the first known DP-SCO algorithm with near optimal privacy-utility whose DP verification scales better than training cost, significantly reducing verification costs on large datasets.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。