arXiv:2410.00542cs.LGcs.CR2024-10中稿 · publication in the…被引 1

提出隐私保护下的主动学习新方法,解决数据选择与隐私的平衡难题。

Differentially Private Active Learning: Balancing Effective Data Selection and Privacy

  • 通过步骤放大机制提升数据参与训练效率
  • 实验证明在视觉和NLP任务上可提升模型性能
  • 适合关注隐私与数据效率平衡的研究者

主动学习(AL)通过迭代选择、标注和训练最有信息量的数据来优化机器学习中的数据标注。然而,其与正式隐私保护方法(尤其是差分隐私,DP)的结合仍鲜有研究。尽管部分工作探索了在线学习场景下的差分隐私主动学习,但在标准学习设置下如何融合两者仍无解,严重限制了主动学习在隐私敏感领域的应用。本文针对标准学习场景提出差分隐私主动学习(DP-AL)。我们发现,将DP-SGD训练直接集成到主动学习中会导致隐私预算分配困难和数据利用率低。为此,我们提出步骤放大策略,利用批次创建中的个体采样概率最大化数据点参与训练的次数,从而优化数据利用。此外,我们评估了多种数据选择获取函数在隐私约束下的有效性,发现许多常用函数在隐私条件下变得不适用。实验表明,DP-AL在视觉和自然语言处理任务中对特定数据集和模型架构能提升性能。但结果也揭示了隐私约束环境下主动学习的局限性,强调了隐私、模型准确率与数据选择准确性之间的权衡。

原文摘要 · Abstract (English)

Active learning (AL) is a widely used technique for optimizing data labeling in machine learning by iteratively selecting, labeling, and training on the most informative data. However, its integration with formal privacy-preserving methods, particularly differential privacy (DP), remains largely underexplored. While some works have explored differentially private AL for specialized scenarios like online learning, the fundamental challenge of combining AL with DP in standard learning settings has remained unaddressed, severely limiting AL's applicability in privacy-sensitive domains. This work addresses this gap by introducing differentially private active learning (DP-AL) for standard learning settings. We demonstrate that naively integrating DP-SGD training into AL presents substantial challenges in privacy budget allocation and data utilization. To overcome these challenges, we propose step amplification, which leverages individual sampling probabilities in batch creation to maximize data point participation in training steps, thus optimizing data utilization. Additionally, we investigate the effectiveness of various acquisition functions for data selection under privacy constraints, revealing that many commonly used functions become impractical. Our experiments on vision and natural language processing tasks show that DP-AL can improve performance for specific datasets and model architectures. However, our findings also highlight the limitations of AL in privacy-constrained environments, emphasizing the trade-offs between privacy, model accuracy, and data selection accuracy.

主动学习差分隐私数据效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。