通过被动攻击检测二分类模型中的标签记忆现象
BLIA: Detect model memorization in binary classification model through passive Label Inference attack
- 利用置信度与损失值设计被动标签推断攻击
- 即使启用差分隐私,成功率仍超50%以上
- 适用于评估模型隐私泄露风险的从业者
模型记忆对机器学习模型的泛化能力及训练数据隐私具有重要影响。本文通过两种新型被动标签推断攻击(BLIA)研究二分类模型中的标签记忆现象。这些攻击无需交互或修改训练过程,仅依赖预训练模型的输出,如置信度分数和对数损失值。通过在受控子集(称为“哨兵”)中故意翻转50%的标签,评估在无标签差分隐私(Label-DP)和基于随机响应的Label-DP两种条件下,标签记忆的程度。尽管采用了不同程度的Label-DP,所提攻击均持续取得超过50%的成功率,显著高于随机猜测基准,明确证明了模型即便在标签与特征刻意无关的情况下仍会记忆训练标签。
原文摘要 · Abstract (English)
Model memorization has implications for both the generalization capacity of machine learning models and the privacy of their training data. This paper investigates label memorization in binary classification models through two novel passive label inference attacks (BLIA). These attacks operate passively, relying solely on the outputs of pre-trained models, such as confidence scores and log-loss values, without interacting with or modifying the training process. By intentionally flipping 50% of the labels in controlled subsets, termed "canaries," we evaluate the extent of label memorization under two conditions: models trained without label differential privacy (Label-DP) and those trained with randomized response-based Label-DP. Despite the application of varying degrees of Label-DP, the proposed attacks consistently achieve success rates exceeding 50%, surpassing the baseline of random guessing and conclusively demonstrating that models memorize training labels, even when these labels are deliberately uncorrelated with the features.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。