厘清了差分隐私优化中关键超参数的真正影响,提升私密学习性能。
R+R:Understanding Hyperparameter Effects in DP-SGD
- 通过因子实验独立验证超参数作用,避免传统研究的干扰。
- 确认梯度裁剪阈值与学习率组合对性能影响最显著。
- 在多个数据集和模型上复现结果,增强结论可信度。
DP-SGD 的关键超参数影响研究缺乏共识、可验证性和可复现性,矛盾甚至轶事性的说法加剧了这一问题。尽管 DP-SGD 是私密机器学习的标准优化算法,但其性能常低于非私密方法,限制了实际应用。合适的超参数设置可改善隐私-效用权衡,理解其影响有助于简化优化过程,推动私密学习采纳。为此,本文开展复现研究:将现有研究归纳为假设,设计专门的因子实验以独立识别超参数效应,并评估这些假设在多个数据集、模型架构和差分隐私预算下的可复现性。结果表明,批量大小与训练轮数的主效应及交互效应无法一致复现;但梯度裁剪阈值与学习率之间的关系可成功复现,且其组合的重要性远高于其他超参数。
原文摘要 · Abstract (English)
Research on the effects of essential hyperparameters of DP-SGD lacks consensus, verification, and replication. Contradictory and anecdotal statements on their influence make matters worse. While DP-SGD is the standard optimization algorithm for privacy-preserving machine learning, its adoption is still commonly challenged by low performance compared to non-private learning approaches. As proper hyperparameter settings can improve the privacy-utility trade-off, understanding the influence of the hyperparameters promises to simplify their optimization towards better performance, and likely foster acceptance of private learning. To shed more light on these influences, we conduct a replication study: We synthesize extant research on hyperparameter influences of DP-SGD into conjectures, conduct a dedicated factorial study to independently identify hyperparameter effects, and assess which conjectures can be replicated across multiple datasets, model architectures, and differential privacy budgets. While we cannot (consistently) replicate conjectures about the main and interaction effects of the batch size and the number of epochs, we were able to replicate the conjectured relationship between the clipping threshold and learning rate. Furthermore, we were able to quantify the significant importance of their combination compared to the other hyperparameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。