arXiv:2601.03429cs.CRcs.LG2026-01中稿 · publication at the…

提出可检测并降低模型解释泄露隐私的系统,保障高风险场景下的安全透明。

DeepLeak: Privacy Enhancing Hardening of Model Explanations Against Membership Leakage

  • 开发新型成员推断攻击,量化解释方法在默认设置下的隐私泄露程度。
  • 通过噪声、裁剪、掩码等轻量策略,减少95%以上泄露,仅损失3.3%解释效果。
  • 揭示稀疏性与敏感度是导致泄露的关键原因,为实践提供指导。

机器学习可解释性对医疗诊断、贷款审批等高风险场景的算法透明至关重要,但这些领域也要求严格隐私保护,导致可解释性与隐私之间存在矛盾。尽管已有研究指出解释方法可能泄露成员信息,但从业者仍缺乏系统性指导来选择或部署兼顾透明与隐私的解释技术。本文提出 DeepLeak 系统,用于审计和缓解后验解释方法的隐私风险。DeepLeak 在三方面推进了当前技术:(1)全面的泄露分析:设计更强的解释感知成员推断攻击(MIA),量化主流解释方法在默认配置下泄露的会员信息量;(2)轻量级加固策略:引入模型无关的实用缓解方法,包括灵敏度校准噪声、归因裁剪和掩码,显著降低会员泄露,同时保持解释效用;(3)根因分析:通过受控实验,识别出驱动泄露的算法特性(如归因稀疏性和敏感度)。在图像基准上评估15种解释技术,结果显示默认设置下泄露量比以往报告高出最多74.9%。所提缓解措施可使泄露降低最高达95%(最低46.5%),平均解释效用损失不超过3.3%。DeepLeak 提供了一条系统化、可复现的安全可解释路径。

原文摘要 · Abstract (English)

Machine learning (ML) explainability is central to algorithmic transparency in high-stakes settings such as predictive diagnostics and loan approval. However, these same domains require rigorous privacy guaranties, creating tension between interpretability and privacy. Although prior work has shown that explanation methods can leak membership information, practitioners still lack systematic guidance on selecting or deploying explanation techniques that balance transparency with privacy. We present DeepLeak, a system to audit and mitigate privacy risks in post-hoc explanation methods. DeepLeak advances the state-of-the-art in three ways: (1) comprehensive leakage profiling: we develop a stronger explanation-aware membership inference attack (MIA) to quantify how much representative explanation methods leak membership information under default configurations; (2) lightweight hardening strategies: we introduce practical, model-agnostic mitigations, including sensitivity-calibrated noise, attribution clipping, and masking, that substantially reduce membership leakage while preserving explanation utility; and (3) root-cause analysis: through controlled experiments, we pinpoint algorithmic properties (e.g., attribution sparsity and sensitivity) that drive leakage. Evaluating 15 explanation techniques across four families on image benchmarks, DeepLeak shows that default settings can leak up to 74.9% more membership information than previously reported. Our mitigations cut leakage by up to 95% (minimum 46.5%) with only <=3.3% utility loss on average. DeepLeak offers a systematic, reproducible path to safer explainability in privacy-sensitive ML.

隐私保护模型解释成员推断

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。