无需参考模型,通过损失分布尾部特征估算模型隐私漏洞。
The Tail Tells All: Estimating Model-Level Membership Inference Vulnerability Without Reference Models
- 利用训练/测试损失分布的尾部缺失现象预测隐私风险
- 在多种架构和数据集上准确估计对最新攻击(LiRA)的脆弱性
- 特别适合评估大语言模型,且比传统方法更高效
会员推理攻击(MIAs)已成为评估人工智能模型隐私风险的标准工具。然而,现有最先进的攻击方法需要训练大量计算成本高昂的参考模型,限制了其实际应用。本文提出一种新方法,可在不依赖参考模型的情况下,估计模型层面的隐私脆弱性——即低误报率下的真阳性率(TPR)。实证分析表明,损失分布具有非对称性和重尾特性,且多数面临攻击风险的样本在训练后已从高损失区域(尾部)移动至低损失区域(头部)。基于此发现,我们仅通过训练和测试损失分布即可构建评估方法:以高损失区域中无异常点作为风险预测指标。我们在多种架构与数据集上评估该方法(即简单损失攻击的TNR),结果表明其能准确估计模型对最先进攻击(LiRA)的脆弱性。同时,该方法优于低成本攻击(如RMIA)及其他分布差异度量。最后,我们探索使用非线性函数评估风险,结果显示该策略在大型语言模型中前景广阔。
原文摘要 · Abstract (English)
Membership inference attacks (MIAs) have emerged as the standard tool for evaluating the privacy risks of AI models. However, state-of-the-art attacks require training numerous, often computationally expensive, reference models, limiting their practicality. We present a novel approach for estimating model-level vulnerability, the TPR at low FPR, to membership inference attacks without requiring reference models. Empirical analysis shows loss distributions to be asymmetric and heavy-tailed and suggests that most points at risk from MIAs have moved from the tail (high-loss region) to the head (low-loss region) of the distribution after training. We leverage this insight to propose a method to estimate model-level vulnerability from the training and testing distribution alone: using the absence of outliers from the high-loss region as a predictor of the risk. We evaluate our method, the TNR of a simple loss attack, across a wide range of architectures and datasets and show it to accurately estimate model-level vulnerability to the SOTA MIA attack (LiRA). We also show our method to outperform both low-cost (few reference models) attacks such as RMIA and other measures of distribution difference. We finally evaluate the use of non-linear functions to evaluate risk and show the approach to be promising to evaluate the risk in large-language models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。