arXiv:2603.04881cs.LGcs.CY2026-03

差分隐私训练损害模型公平性与鲁棒性,根源在于特征噪声比下降

Differential Privacy in Two-Layer Networks: How DP-SGD Harms Fairness and Robustness

  • 从特征学习角度构建统一分析框架,揭示隐私噪声对特征提取的抑制机制
  • 隐私噪声导致类别间特征噪声比失衡,引发不公平;对长尾语义数据影响更严重
  • 隐私训练加剧对抗攻击脆弱性,且预训练-微调范式在分布偏移下无效

差分隐私学习对敏感数据建模至关重要,但实证研究显示其会降低性能、引发公平性问题(如差别影响)并削弱对抗鲁棒性。本文针对两层ReLU卷积神经网络,提出统一的特征中心分析框架,研究差分隐私随机梯度下降(DP-SGD)的特征学习动态。理论证明测试损失受关键指标‘特征-噪声比’(FNR)约束。研究表明:1)类别与子群体间不平衡的FNR导致差别影响;2)同一类别中,语义长尾数据受噪声负面影响更大;3)噪声注入加剧对抗攻击脆弱性。此外,主流的公开预训练+私有微调范式在数据分布显著偏移时亦无法保证性能提升。合成与真实数据实验验证了理论结论。

原文摘要 · Abstract (English)

Differentially private learning is essential for training models on sensitive data, but empirical studies consistently show that it can degrade performance, introduce fairness issues like disparate impact, and reduce adversarial robustness. The theoretical underpinnings of these phenomena in modern, non-convex neural networks remain largely unexplored. This paper introduces a unified feature-centric framework to analyze the feature learning dynamics of differentially private stochastic gradient descent (DP-SGD) in two-layer ReLU convolutional neural networks. Our analysis establishes test loss bounds governed by a crucial metric: the feature-to-noise ratio (FNR). We demonstrate that the noise required for privacy leads to suboptimal feature learning, and specifically show that: 1) imbalanced FNRs across classes and subpopulations cause disparate impact; 2) even in the same class, noise has a greater negative impact on semantically long-tailed data; and 3) noise injection exacerbates vulnerability to adversarial attacks. Furthermore, our analysis reveals that the popular paradigm of public pre-training and private fine-tuning does not guarantee improvement, particularly under significant feature distribution shifts between datasets. Experiments on synthetic and real-world data corroborate our theoretical findings.

差分隐私公平性鲁棒性特征学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。