探究可解释性、隐私与模型性能在解释辅助模型提取攻击下的权衡关系。
On the interplay of Explainability, Privacy and Predictive Performance with Explanation-assisted Model Extraction
- 对比训练阶段与解释阶段应用差分隐私的两种策略。
- 发现解释阶段加噪能更好保护隐私且对性能影响更小。
- 适合关注模型安全与可解释性平衡的研究者参考。
机器学习即服务(MLaaS)因其便捷性受到广泛关注,使组织无需大量投入即可使用先进分析工具。然而,此类平台面临模型提取攻击(MEA)等安全与隐私威胁。随着可解释人工智能(XAI)在MLaaS中的集成,攻击者可能利用反事实解释(CFs)来促进模型提取。本文研究了在采用差分隐私(DP)以缓解由反事实解释引发的模型提取攻击时,模型性能、隐私与可解释性之间的权衡关系。我们评估了两种不同的DP策略:一种在分类模型训练过程中实施,另一种在生成反事实解释时应用。
原文摘要 · Abstract (English)
Machine Learning as a Service (MLaaS) has gained important attraction as a means for deploying powerful predictive models, offering ease of use that enables organizations to leverage advanced analytics without substantial investments in specialized infrastructure or expertise. However, MLaaS platforms must be safeguarded against security and privacy attacks, such as model extraction (MEA) attacks. The increasing integration of explainable AI (XAI) within MLaaS has introduced an additional privacy challenge, as attackers can exploit model explanations particularly counterfactual explanations (CFs) to facilitate MEA. In this paper, we investigate the trade offs among model performance, privacy, and explainability when employing Differential Privacy (DP), a promising technique for mitigating CF facilitated MEA. We evaluate two distinct DP strategies: implemented during the classification model training and at the explainer during CF generation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。