用合成数据攻击方法检测反事实样本隐私风险,发现无需模型访问也能泄露训练数据。
Quantifying the Privacy of Counterfactuals by Leveraging Membership Inference Attacks Against Synthetic Data

- 借鉴合成数据的成员推断攻击,评估反事实样本的隐私暴露程度
- 仅凭反事实数据集即可成功发起成员推断攻击,无需查询目标模型
- 提醒模型开发者:发布反事实解释可能引发隐私泄露,需谨慎处理
反事实样本常用于高风险决策场景,通过展示用户特征变化如何导致期望结果来解释机器学习模型。然而,利用反事实解释也可能被攻击者用于针对模型或其训练数据的隐私攻击。本文基于反事实与合成数据的相似性,证明可沿用针对合成数据设计的成员推断攻击来威胁反事实隐私。具体而言,我们评估了现有合成数据成员推断攻击在多种反事实生成方法上的有效性。值得注意的是,现有针对反事实的攻击通常需要访问目标模型,而本文展示即使仅有反事实数据集本身,无需模型查询,仍可成功实施成员推断攻击。结果表明,模型开发者在向用户释放反事实解释时应更加谨慎,以免造成隐私泄露。
原文摘要 · Abstract (English)
Counterfactuals are typically used in high-stakes decision areas to explain a machine learning model by showing how changes to the user profiles result in the desired outcome. However, explaining the model's decisions through counterfactuals can also be exploited by an adversary to conduct privacy attacks against the model or its training data. Drawing on the analogy that counterfactuals provide realistic substitutes for real training data, similar to synthetic data, we demonstrate in this paper how it is possible to successfully perform privacy attacks on counterfactuals by drawing on the attacks developed against synthetic data. More precisely, we investigate the effectiveness of the membership inference attacks designed for synthetic data on various types of counterfactuals. Additionally, while existing membership inference attacks against counterfactuals usually require to be able to query the model, we show how it is possible to perform successful membership inference attacks using only a set of counterfactuals, with no access to the model from which they are generated. Our results demonstrate that model developers should be more cautious when releasing counterfactuals to various users, as it can lead to a privacy breach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。