arXiv:2511.10284cs.AI2025-11中稿 · AAAI

用反事实推理识别模型决策中的隐私泄露,让审计更透明可信。

Beyond Verification: Abductive Explanations for Post-AI Assessment of Privacy Leakage

  • 通过反事实解释找出支撑模型决策的最小证据集
  • 在德国信贷数据集上验证敏感特征对隐私泄露的影响
  • 既保证隐私可审计,又生成人能理解的解释,适合监管与合规场景

AI决策过程中的隐私泄露带来重大风险,尤其当敏感信息可被推断时。本文提出一种形式化框架,利用反事实解释来审计隐私泄露,识别出支撑模型决策的最小充分证据,并判断是否披露了敏感信息。该框架同时形式化个体与系统层面的泄露,引入潜在适用解释(PAE)概念,用于识别其结果可掩盖敏感特征个体的群体。该方法在德国信贷数据集上的实验表明,敏感属性在模型决策中的重要性直接影响隐私泄露程度。尽管存在计算复杂性和简化假设,研究结果证明反事实推理可实现可解释的隐私审计,为人工智能决策中的透明性、可解释性与隐私保护提供可行路径。

原文摘要 · Abstract (English)

Privacy leakage in AI-based decision processes poses significant risks, particularly when sensitive information can be inferred. We propose a formal framework to audit privacy leakage using abductive explanations, which identifies minimal sufficient evidence justifying model decisions and determines whether sensitive information disclosed. Our framework formalizes both individual and system-level leakage, introducing the notion of Potentially Applicable Explanations (PAE) to identify individuals whose outcomes can shield those with sensitive features. This approach provides rigorous privacy guarantees while producing human understandable explanations, a key requirement for auditing tools. Experimental evaluation on the German Credit Dataset illustrates how the importance of sensitive literal in the model decision process affects privacy leakage. Despite computational challenges and simplifying assumptions, our results demonstrate that abductive reasoning enables interpretable privacy auditing, offering a practical pathway to reconcile transparency, model interpretability, and privacy preserving in AI decision-making.

隐私审计反事实解释可解释AI

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。