arXiv:2605.29454cs.LG2026-05

提出全流程隐私评估框架,系统分析机器学习中成员推断攻击的实效性。

A Full-Pipeline Framework for Evaluating Membership Inference Attacks in Machine Learning

论文配图:A Full-Pipeline Framework for Evaluating Membership Inference Attacks in Machine Learning
图 1 · 摘自论文原文
  • 构建覆盖数据、模型、算法全链路的评估体系,支持多场景测试。
  • 实证发现攻击效果高度依赖威胁模型与评价指标选择。
  • 提供可直接使用的审计工具包,助力实际隐私风险研判。

尽管成员推断攻击(MIA)是识别训练数据的主要方法,其应用已扩展至隐私审计和机器遗忘等领域,但该领域缺乏系统框架来评估不同上下文对MIA有效性的影响。缺乏这种刻画会使实践者误将基准表现良好的算法部署于真实数据时面临统计失效风险。为此,我们提出一个全面的评估框架,系统刻画从数据到后处理模块的全链条隐私风险。框架涵盖多种训练配置,并引入三种互补指标:平衡准确率用于对称成本场景,低FPR下的真正率(TPR)或低FNR下的真负率(TNR)用于不对称成本场景(如严控误报或漏检)。针对现有MIA假设对手能力差异,我们形式化两种标准化威胁模型,并相应适配攻击变体以实现公平比较。大量实证表明,特定MIA方法的效力显著受威胁模型和评估指标影响。最终,我们提炼出可操作建议并提供即用型审计工具包,帮助实践者开展更可靠的隐私评估。

原文摘要 · Abstract (English)

While Membership Inference Attacks (MIAs) are the prevailing method for identifying training data, their application has expanded into privacy auditing and machine unlearning. Nevertheless, the field lacks a systematic framework for evaluating how different contexts affect MIA efficacy. Without such a characterization, practitioners risk deploying algorithms that perform well on benchmarks but become statistically irrelevant when faced with the nuances of specific, real-world datasets. To bridge this gap and provide actionable insights, we introduce a comprehensive evaluation framework that systematically characterizes privacy risks across the entire machine learning pipeline, spanning data, architectures, algorithms, and post-training modules. Designed to inherently capture diverse operational contexts, our framework rigorously evaluates state-of-the-art MIAs across a broad spectrum of training configurations. To account for varying misclassification costs in real-world deployments, we employ three complementary metrics: Balanced Accuracy for symmetric costs, alongside TPR at low FPR (or TNR at low FNR) for asymmetric scenarios where false alarms or missed detections are strictly penalized. Furthermore, recognizing that existing MIAs assume divergent adversary capabilities, we formalize two standardized threat models and adapt these attacks into corresponding variants to ensure an equitable benchmark. Extensive empirical evaluations demonstrate that the efficacy of specific MIA methodologies is highly sensitive to the assumed threat models and chosen evaluation metrics. Ultimately, we distill these findings into actionable guidelines and provide a ready-to-use auditing toolkit, empowering practitioners to conduct better privacy assessments.

隐私审计成员推断评估框架

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。