提出可控制误发现率的成员推理攻击方法,提升攻击可信度。
Membership Inference Attacks with False Discovery Rate Control
- 设计新攻击方法,实现对误发现率的严格控制
- 在黑盒与持续学习场景下均验证有效性能
- 可作为后处理模块无缝集成现有攻击方法
深度学习模型易受成员推理攻击(MIAs)威胁,旨在判断数据记录是否用于训练目标模型。尽管MIAs重要且广泛应用,但现有方法缺乏对误发现率(FDR)的保障——即被识别为正例中虚假发现的期望比例。由于底层分布未知且非成员概率估计存在依赖性,实现FDR保障极具挑战。本文提出一种新型成员推理攻击方法,可提供严格的FDR控制,并同时保证标记真实非成员为成员的边际概率。该方法可作为后处理封装器,无缝集成至现有MIA方法中,适用于黑盒设置与持续学习等多样场景。理论分析与大量实验验证了其优越性能。
原文摘要 · Abstract (English)
Recent studies have shown that deep learning models are vulnerable to membership inference attacks (MIAs), which aim to infer whether a data record was used to train a target model or not. To analyze and study these vulnerabilities, various MIA methods have been proposed. Despite the significance and popularity of MIAs, existing works on MIAs are limited in providing guarantees on the false discovery rate (FDR), which refers to the expected proportion of false discoveries among the identified positive discoveries. However, it is very challenging to ensure the false discovery rate guarantees, because the underlying distribution is usually unknown, and the estimated non-member probabilities often exhibit interdependence. To tackle the above challenges, in this paper, we design a novel membership inference attack method, which can provide the guarantees on the false discovery rate. Additionally, we show that our method can also provide the marginal probability guarantee on labeling true non-member data as member data. Notably, our method can work as a wrapper that can be seamlessly integrated with existing MIA methods in a post-hoc manner, while also providing the FDR control. We perform the theoretical analysis for our method. Extensive experiments in various settings (e.g., the black-box setting and the lifelong learning setting) are also conducted to verify the desirable performance of our method.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。