用户查询反事实解释时,保护其特征隐私并精准获取最邻近反事实实例。
What If, But Privately: Private Counterfactual Retrieval
- 基于信息论设计隐私保护机制,确保用户特征不泄露。
- 在保持用户隐私的前提下,从数据库中精确检索最近邻反事实样本。
- 支持不可变特征与用户偏好,适用于医疗、金融等高敏感场景。
在高风险应用中,黑箱机器学习模型的透明性与可解释性至关重要。提供反事实解释是满足这一需求的一种方式,但可能威胁到提供解释机构及请求用户的隐私。本文关注用户隐私:在不向机构暴露其特征向量的前提下,检索反事实实例。提出私有反事实检索(PCR)框架,在信息论上实现用户隐私的完美保护,并可精确检索数据库中最近邻的反事实解释。进一步提出两种方案,减少对机构数据库的信息泄露。针对部分特征不可变的情况,提出不可变私有反事实检索(I-PCR)方案,保护用户特征与不可变集隐私。还引入用户偏好以生成更具行动力的解释。数值实验验证了理论结果,并比较了各方案的数据库泄露程度。
原文摘要 · Abstract (English)
Transparency and explainability are two important aspects to be considered when employing black-box machine learning models in high-stake applications. Providing counterfactual explanations is one way of catering this requirement. However, this also poses a threat to the privacy of the institution that is providing the explanation, as well as the user who is requesting it. In this work, we are primarily concerned with the user's privacy who wants to retrieve a counterfactual instance, without revealing their feature vector to the institution. Our framework retrieves the exact nearest neighbor counterfactual explanation from a database of accepted points while achieving perfect, information-theoretic, privacy for the user. First, we introduce the problem of private counterfactual retrieval (PCR) and propose a baseline PCR scheme that keeps the user's feature vector information-theoretically private from the institution. Building on this, we propose two other schemes that reduce the amount of information leaked about the institution database to the user, compared to the baseline scheme. Second, we relax the assumption of mutability of all features, and consider the setting of immutable PCR (I-PCR). Here, the user retrieves the nearest counterfactual without altering a private subset of their features, which constitutes the immutable set, while keeping their feature vector and immutable set private from the institution. For this, we propose two schemes that preserve the user's privacy information-theoretically, but ensure varying degrees of database privacy. Third, we extend our PCR and I-PCR schemes to incorporate user's preference on transforming their attributes, so that a more actionable explanation can be received. Finally, we present numerical results to support our theoretical findings, and compare the database leakage of the proposed schemes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。