不重训练模型也能精准评估单个数据的隐私泄露风险。
Assessing Per-Sample Membership Inference Vulnerability without Retraining
- 用损失和几何特征联合衡量单样本隐私暴露度
- 在多种模型上识别出最高风险样本,效果优于传统方法
- 仅需一个训练好的模型,无需构建影子模型
近期研究表明,针对特定样本的成员推理攻击(MIAs)性能远超无目标攻击。本文提出:能否在不训练影子模型的前提下评估单个训练样本的隐私脆弱性?研究发现,样本的隐私暴露不仅取决于其损失值,还受数据相关的几何特性影响。在线性设置下,我们推导出黑箱攻击下个体隐私脆弱性的闭式分解,由群体杠杆率与残差损失项构成,明确揭示了样本几何如何转化为隐私风险。由于多数现代架构的最终层为线性结构,我们将该框架推广至深度网络,提出一种基于最后层表示的代理评分方法,仅需一个训练好的模型,无需影子模型。在多种数据集和架构上的实证评估表明,该评分在识别高风险样本方面优于损失和梯度范数基线,在先进攻击下表现更优,提供了一种计算高效且理论扎实的单样本隐私风险评估工具。
原文摘要 · Abstract (English)
Recent work in the privacy literature shows that sample-targeted membership inference attacks (MIAs) significantly outperform untargeted approaches by a wide margin. Motivated by this observation, we address the following question: can the privacy vulnerability of individual training points be assessed without training shadow models? We show that per-sample exposure to MIA is governed not only by a point's loss, but also by a data-dependent geometric measure. In the linear setting, we derive a closed-form decomposition of individual black-box MIA vulnerability into a population leverage score and a residual loss term, making explicit how sample-dependent geometry translates into privacy exposure. Since the final layer of most modern architectures is linear, we extend this framework to deep networks and propose a surrogate score operating on last-layer representations that requires only a single trained model and no shadow models. Empirical evaluations across diverse datasets and architectures show that our score outperforms loss and gradient-norm baselines at identifying the highest-risk points under state-of-the-art attacks, providing a computationally efficient and theoretically grounded tool for per-sample privacy risk assessment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。