研究隐私机器学习中公私数据分布偏移的影响,发现小偏移时需充足数据,大偏移时公有数据无用。
Lower Bounds for Public-Private Learning under Distribution Shift
- 分析公私数据在分布偏移下的联合学习下界
- 小偏移时需充足公/私数据才能准确估计参数
- 大偏移时公有数据无法提升性能,适合关注隐私学习的学者
目前最有效的差分隐私机器学习算法依赖额外的所谓公开数据源。当两个数据源结合时若能产生协同效应,则该范式最具吸引力。然而在均值估计等场景中已有强下界表明:当两数据源分布相同时,联合使用并无额外价值。本文将已知的公私学习下界扩展至两数据源存在显著分布偏移的情形。结果适用于高斯均值估计(两分布均值不同)和高斯线性回归(参数偏移)。发现当偏移较小时(相对于目标精度),公有或私有数据必须足够丰富才能估计私有参数;反之,当偏移较大时,公有数据无法带来任何增益。
原文摘要 · Abstract (English)
The most effective differentially private machine learning algorithms in practice rely on an additional source of purportedly public data. This paradigm is most interesting when the two sources combine to be more than the sum of their parts. However, there are settings such as mean estimation where we have strong lower bounds, showing that when the two data sources have the same distribution, there is no complementary value to combining the two data sources. In this work we extend the known lower bounds for public-private learning to setting where the two data sources exhibit significant distribution shift. Our results apply to both Gaussian mean estimation where the two distributions have different means, and to Gaussian linear regression where the two distributions exhibit parameter shift. We find that when the shift is small (relative to the desired accuracy), either public or private data must be sufficiently abundant to estimate the private parameter. Conversely, when the shift is large, public data provides no benefit.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。