arXiv:2502.14309cs.LGcs.IT2025-02

研究标签差分隐私下的学习极限,发现仅保护标签能显著提升性能。

On Theoretical Limits of Learning with Label Differential Privacy

  • 将任务转化为多假设检验,推导理论下界
  • 本地标签DP收敛速度远快于全量DP
  • 适合关注隐私保护与模型性能平衡的研究者

标签差分隐私(Label DP)针对标签私有、特征公开的学习场景设计。尽管已有多种方法在标签DP下进行学习,但其理论极限仍不明确。本文在本地和中心化模型下,对分类与回归任务的标签DP学习进行了基础性分析,以极小极大收敛率刻画其理论边界。通过将各任务转化为多重假设检验问题并界定检验误差,建立下界;同时设计算法实现匹配的上界。结果表明:在标签本地差分隐私(LDP)下,风险收敛速度远快于全量差分隐私(即同时保护特征与标签),说明仅保护标签的放松定义具有显著优势;而在标签中心化差分隐私(CDP)下,性能仅比全量DP提升常数倍,表明该放松带来的收益有限。

原文摘要 · Abstract (English)

Label differential privacy (DP) is designed for learning problems involving private labels and public features. While various methods have been proposed for learning under label DP, the theoretical limits remain largely unexplored. In this paper, we investigate the fundamental limits of learning with label DP in both local and central models for both classification and regression tasks, characterized by minimax convergence rates. We establish lower bounds by converting each task into a multiple hypothesis testing problem and bounding the test error. Additionally, we develop algorithms that yield matching upper bounds. Our results demonstrate that under label local DP (LDP), the risk has a significantly faster convergence rate than that under full LDP, i.e. protecting both features and labels, indicating the advantages of relaxing the DP definition to focus solely on labels. In contrast, under the label central DP (CDP), the risk is only reduced by a constant factor compared to full DP, indicating that the relaxation of CDP only has limited benefits on the performance.

差分隐私理论分析机器学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。