arXiv:2503.06732cs.LG2025-03

提升私有训练的数据效率,改进模型收敛速度

Data Efficient Subset Training with Differential Privacy

  • 将高效数据训练方法GLISTER适配到私有训练场景
  • 发现实际隐私预算对数据高效训练限制过强
  • 为私有学习中的高效训练提供新思路

私有机器学习在隐私预算与训练性能之间存在权衡。私有训练的收敛速度显著变慢,且需要大量超参数调优。因此,如何高效地进行私有模型训练成为研究重点。本文研究数据高效训练方法在私有训练环境下的有效性,将GLISTER(Killamsetty et al., 2021b)方法适配至私有设置,并进行了广泛评估。实验发现,在实际应用中设定的隐私预算对数据高效训练而言过于严格,限制了其潜力发挥。

原文摘要 · Abstract (English)

Private machine learning introduces a trade-off between the privacy budget and training performance. Training convergence is substantially slower and extensive hyper parameter tuning is required. Consequently, efficient methods to conduct private training of models is thoroughly investigated in the literature. To this end, we investigate the strength of the data efficient model training methods in the private training setting. We adapt GLISTER (Killamsetty et al., 2021b) to the private setting and extensively assess its performance. We empirically find that practical choices of privacy budgets are too restrictive for data efficient training in the private setting.

私有训练数据效率模型优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。