用隐私保护技术降低可解释AI泄露个人数据的风险
Towards integration of Privacy Enhancing Technologies in Explainable Artificial Intelligence
- 引入隐私增强技术防御特征型可解释AI的隐私泄露
- 最佳情况下隐私攻击风险降低49.47%且保持模型性能
- 适合关注隐私安全的AI系统部署与合规团队
可解释人工智能(XAI)虽能缓解黑箱AI决策不透明问题,但其解释过程可能泄露训练或查询数据中个体的隐私。已有研究证明,攻击者可通过解释推断敏感个人信息。当前在生产环境和机器学习即服务系统中,尚缺乏针对此类隐私攻击的有效防御。本文探索隐私增强技术(PETs)作为应对特征型XAI解释中属性推断攻击的防御手段。我们实证评估了三种PETs——合成训练数据、差分隐私训练和噪声添加——在两类特征型XAI上的表现。结果表明,不同方法对隐私保护效果和系统其他属性(如模型效用与解释质量)的影响各异。在最优情况下,集成PETs使攻击风险降低49.47%,同时维持模型效用与解释质量。研究还提出了在XAI中最大化收益、最小化隐私攻击风险的策略。
原文摘要 · Abstract (English)
Explainable Artificial Intelligence (XAI) is a crucial pathway in mitigating the risk of non-transparency in the decision-making process of black-box Artificial Intelligence (AI) systems. However, despite the benefits, XAI methods are found to leak the privacy of individuals whose data is used in training or querying the models. Researchers have demonstrated privacy attacks that exploit explanations to infer sensitive personal information of individuals. Currently there is a lack of defenses against known privacy attacks targeting explanations when vulnerable XAI are used in production and machine learning as a service system. To address this gap, in this article, we explore Privacy Enhancing Technologies (PETs) as a defense mechanism against attribute inference on explanations provided by feature-based XAI methods. We empirically evaluate 3 types of PETs, namely synthetic training data, differentially private training and noise addition, on two categories of feature-based XAI. Our evaluation determines different responses from the mitigation methods and side-effects of PETs on other system properties such as utility and performance. In the best case, PETs integration in explanations reduced the risk of the attack by 49.47%, while maintaining model utility and explanation quality. Through our evaluation, we identify strategies for using PETs in XAI for maximizing benefits and minimizing the success of this privacy attack on sensitive personal information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。