arXiv:2503.21528stat.MLcs.CR2025-03

提出新型差分隐私机制,提升小样本和不平衡数据下的模型性能。

Bayesian Pseudo Posterior Mechanism for Differentially Private Machine Learning

  • 用伪后验分布按记录泄露风险加权,降低高风险数据影响。
  • 在相同隐私预算下,比标准DP-SGD准确率更高,仅轻微损失实用性能。
  • 适用于数据不平衡或标注样本少的现实场景,如政府统计建模。

差分隐私(DP)在部署机器学习应用中日益重要,因其能为训练数据中个人隐私提供强保障。然而,当前常用的DP机制在真实世界分布(如高度不平衡或小规模标注数据集)上表现不佳。本文提出一种可扩展的深度学习新DP机制SWAG-PPM,通过伪后验分布按记录的披露风险成比例地降低其似然贡献作为随机化机制。以美国职业安全与健康管理局(OSHA)发布的高度不平衡公开数据集为例,我们在工作场所伤害文本分类任务中验证了该方法。结果表明,在相似隐私预算下,SWAG-PPM相较于非私有基线仅有轻微效用下降,但显著优于行业标准的DP-SGD。

原文摘要 · Abstract (English)

Differential privacy (DP) is becoming increasingly important for deployed machine learning applications because it provides strong guarantees for protecting the privacy of individuals whose data is used to train models. However, DP mechanisms commonly used in machine learning tend to struggle on many real world distributions, including highly imbalanced or small labeled training sets. In this work, we propose a new scalable DP mechanism for deep learning models, SWAG-PPM, by using a pseudo posterior distribution that downweights by-record likelihood contributions proportionally to their disclosure risks as the randomized mechanism. As a motivating example from official statistics, we demonstrate SWAG-PPM on a workplace injury text classification task using a highly imbalanced public dataset published by the U.S. Occupational Safety and Health Administration (OSHA). We find that SWAG-PPM exhibits only modest utility degradation against a non-private comparator while greatly outperforming the industry standard DP-SGD for a similar privacy budget.

差分隐私深度学习数据不平衡伪后验

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。