用少量样本实现高效隐私攻击,评估模型训练数据泄露风险。
Membership Inference Attacks fueled by Few-Short Learning to detect privacy leakage tackling data integrity
- 基于少样本学习构建新型隐私攻击模型,降低资源需求。
- 在低误报率下仍保持高检测准确率,量化隐私泄露程度。
- 提出可解释的评估指标,适合安全与合规场景使用。
深度学习模型会记忆部分训练数据,导致隐私泄露。成员推理攻击(MIA)可利用此特性推断某数据是否用于模型训练,进而窃取敏感信息。现有最优攻击方法虽效果显著,但资源要求过高,难以作为实用工具评估隐私风险。此外,主流评价指标在低误报率下的真正例率缺乏可解释性。本文提出双重改进:一是基于少样本学习的FeS-MIA模型,大幅降低评估隐私泄露所需的资源;二是提出可解释的定量与定性评估指标——Log-MIA。实验表明,在图像分类与语言建模任务中,该方法仅需极少额外信息即可有效揭示模型隐私泄露情况,优于现有方法。
原文摘要 · Abstract (English)
Deep learning models have an intrinsic privacy issue as they memorize parts of their training data, creating a privacy leakage. Membership Inference Attacks (MIA) exploit it to obtain confidential information about the data used for training, aiming to steal information. They can be repurposed as a measurement of data integrity by inferring whether it was used to train a machine learning model. While state-of-the-art attacks achieve a significant privacy leakage, their requirements are not feasible enough, hindering their role as practical tools to assess the magnitude of the privacy risk. Moreover, the most appropriate evaluation metric of MIA, the True Positive Rate at low False Positive Rate lacks interpretability. We claim that the incorporation of Few-Shot Learning techniques to the MIA field and a proper qualitative and quantitative privacy evaluation measure should deal with these issues. In this context, our proposal is twofold. We propose a Few-Shot learning based MIA, coined as the FeS-MIA model, which eases the evaluation of the privacy breach of a deep learning model by significantly reducing the number of resources required for the purpose. Furthermore, we propose an interpretable quantitative and qualitative measure of privacy, referred to as Log-MIA measure. Jointly, these proposals provide new tools to assess the privacy leakage and to ease the evaluation of the training data integrity of deep learning models, that is, to analyze the privacy breach of a deep learning model. Experiments carried out with MIA over image classification and language modeling tasks and its comparison to the state-of-the-art show that our proposals excel at reporting the privacy leakage of a deep learning model with little extra information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。