arXiv:2412.02230cs.LG2024-12被引 1

保护隐私的多分类学习,隐藏敏感标签仍能准确分类。

Learning from Concealed Labels

  • 用随机非敏感标签隐藏敏感标签进行标注
  • 在弱假设下建立无偏估计器,收敛率最优
  • 适合医疗、金融等敏感数据场景

在真实场景中,标注敏感标签(如疾病、吸烟)可能威胁个人隐私。为此,我们提出一种新设置——从隐藏标签中学习,用于多分类任务。该方法在标签收集阶段将敏感标签排除在标签集外,仅使用部分随机采样的非敏感标签作为隐藏标签来标注敏感数据。本文在弱假设下建立了从隐藏数据中学习的无偏估计器,所学分类器不仅能准确区分非敏感标签,还能识别敏感标签。我们给出了估计误差的界,并证明分类器达到最优参数收敛率。实验表明,该方法在合成与真实数据集上均有效且显著。

原文摘要 · Abstract (English)

Annotating data for sensitive labels (e.g., disease, smoking) poses a potential threats to individual privacy in many real-world scenarios. To cope with this problem, we propose a novel setting to protect privacy of each instance, namely learning from concealed labels for multi-class classification. Concealed labels prevent sensitive labels from appearing in the label set during the label collection stage, which specifies none and some random sampled insensitive labels as concealed labels set to annotate sensitive data. In this paper, an unbiased estimator can be established from concealed data under mild assumptions, and the learned multi-class classifier can not only classify the instance from insensitive labels accurately but also recognize the instance from the sensitive labels. Moreover, we bound the estimation error and show that the multi-class classifier achieves the optimal parametric convergence rate. Experiments demonstrate the significance and effectiveness of the proposed method for concealed labels in synthetic and real-world datasets.

隐私保护多分类隐藏标签

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。