arXiv:2502.18986cs.CRcs.LG2025-02中稿 · SiMLA workshop 202…被引 1

首次系统评估异构数据下成员推理攻击的性能差异

Evaluating Membership Inference Attacks in heterogeneous-data setups

  • 提出连续度量表衡量表格数据分布的异构性
  • 发现攻击准确率在90%与50%间波动,取决于数据设置
  • 揭示现有研究缺乏统一异构数据模拟基准

在各类机器学习隐私攻击中,成员推理攻击(MIA)最受关注。攻击者通过模型和一个数据点,判断该数据是否参与过训练,且可使用辅助数据集优化攻击算法。现有研究多假设攻击者与目标数据来自同分布,虽便于实验,但现实中罕见。本文首次将MIA评估拓展至异构数据场景:首先设计度量指标,量化任意两组表格数据分布间的异构程度;其次比较两种模拟异构性的方法,结果呈现极端差异——攻击准确率从90%降至50%(即随机猜测)。研究显示,MIA表现高度依赖实验设定;即便已有研究提及异构性,仍缺乏标准化模拟基准。这一缺失严重制约真实场景下的隐私风险评估。

原文摘要 · Abstract (English)

Among all privacy attacks against Machine Learning (ML), membership inference attacks (MIA) attracted the most attention. In these attacks, the attacker is given an ML model and a data point, and they must infer whether the data point was used for training. The attacker also has an auxiliary dataset to tune their inference algorithm. Attack papers commonly simulate setups in which the attacker's and the target's datasets are sampled from the same distribution. This setting is convenient to perform experiments, but it rarely holds in practice. ML literature commonly starts with similar simplifying assumptions (i.e., "i.i.d." datasets), and later generalizes the results to support heterogeneous data distributions. Similarly, our work makes a first step in the generalization of the MIA evaluation to heterogeneous data. First, we design a metric to measure the heterogeneity between any pair of tabular data distributions. This metric provides a continuous scale to analyze the phenomenon. Second, we compare two methodologies to simulate a data heterogeneity between the target and the attacker. These setups provide opposite performances: 90% attack accuracy vs. 50% (i.e., random guessing). Our results show that the MIA accuracy depends on the experimental setup; and even if research on MIA considers heterogeneous data setups, we have no standardized baseline of how to simulate it. The lack of such a baseline for MIA experiments poses a significant challenge to risk assessments in real-world machine learning scenarios.

隐私攻击成员推理异构数据

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。