arXiv:2602.03611cs.LG2026-02

解释生成会泄露隐私,需用差分隐私与主动学习防护

Explanations Leak: Membership Inference with Differential Privacy and Active Learning Defense

  • 将差分隐私与主动学习结合,减少模型记忆和数据暴露
  • 查询式接口暴露反事实解释会增强成员推理攻击效果
  • 在隐私、性能与解释质量间存在三重权衡,需谨慎平衡

反事实解释(CFs)正被越来越多地集成到机器学习即服务(MLaaS)系统中以提升透明度;然而,通过API部署的机器学习模型已面临成员推理攻击(MIA)和模型提取等隐私威胁,而解释机制对这一威胁格局的影响尚未充分理解。本文研究了CFs如何通过查询式API扩大MLaaS的攻击面,从而强化成员推理攻击,并探讨了在不损害模型效用与可解释性的前提下设计防御机制的必要性。首先,我们系统分析了通过查询接口暴露CFs如何增强基于影子模型的成员推理攻击。其次,提出一种融合差分隐私(DP)与主动学习(AL)的防御框架,联合降低模型记忆能力并限制有效训练数据暴露。最后,通过广泛的实证评估,刻画了隐私泄露、预测性能与解释质量之间的三重权衡。研究结果强调,在负责任地部署可解释的MLaaS系统时,必须审慎平衡透明性、效用与隐私。

原文摘要 · Abstract (English)

Counterfactual explanations (CFs) are increasingly integrated into Machine Learning as a Service (MLaaS) systems to improve transparency; however, ML models deployed via APIs are already vulnerable to privacy attacks such as membership inference and model extraction, and the impact of explanations on this threat landscape remains insufficiently understood. In this work, we focus on the problem of how CFs expand the attack surface of MLaaS by strengthening membership inference attacks (MIAs), and on the need to design defense mechanisms that mitigate this emerging risk without undermining utility and explainability. First, we systematically analyze how exposing CFs through query-based APIs enables more effective shadow-based MIAs. Second, we propose a defense framework that integrates Differential Privacy (DP) with Active Learning (AL) to jointly reduce memorization and limit effective training data exposure. Finally, we conduct an extensive empirical evaluation to characterize the three-way trade-off between privacy leakage, predictive performance, and explanation quality. Our findings highlight the need to carefully balance transparency, utility, and privacy in the responsible deployment of explainable MLaaS systems.

隐私保护可解释性差分隐私成员推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。