用梯度引导的解释一致性,防御隐私泄露的成员推断攻击。
TIER: Trajectory-Invariant Explanation Regularization for Membership Privacy

- 通过梯度扰动模拟置信度下降轨迹,用KL散度约束解释分布一致。
- 在多个数据集上使成员推断攻击成功率降至10%以下,同时保持模型精度。
- 适合关注模型可解释性与隐私安全的开发者和研究者使用。
可解释性是构建可信AI的核心,但解释界面可能被攻击者利用,扩大隐私攻击面。近期研究表明,高级成员推断攻击通过利用梯度引导扰动引发的置信度下降轨迹作为判别特征,而非直接使用置信度或解释向量。现有防御方法无法有效应对此类解释驱动的攻击。本文提出一种轨迹不变解释正则化(TIER)防御机制,在训练中利用模型自身梯度作为防御信号,使成员与非成员的解释特征分布趋于一致。该方法通过惩罚梯度引导扰动下的置信度下降波动,并采用KL散度最小化分布偏移。与传统对抗训练侧重标签鲁棒性不同,TIER聚焦解释鲁棒性,通过KL散度实现自一致性约束并降低成员与非成员间置信度下降方差。大量实验表明,该方法能有效缓解攻击,保护隐私的同时维持模型性能与解释保真度。
原文摘要 · Abstract (English)
Explainability is central to building trustworthy AI, yet explanation interfaces can inadvertently provide adversaries with an expanded privacy-related attack surfaces. Recent studies show that advanced membership-inference attacks succeed by exploiting confidence-drop trajectories, induced through attribution-guided perturbations, as discriminative features, rather than directly using confidence scores or explanation vectors. Existing defenses against membership inference fail to directly mitigate such explanation-driven attacks. In this work, we investigate whether, during training, a model's own gradients can be leveraged as defense signals against such attacks, thereby aligning explanation profiles between members and non-members. To this end, we propose a Trajectory-Invariant Explanation Regularization (TIER) defense that penalizes erratic fluctuations in confidence drops simulated through gradient-guided perturbations and simultaneously minimizes the distributional shifts via KL-divergence. Unlike conventional adversarial training, which emphasizes label robustness, our approach targets explanation robustness by enforcing self-consistency through KL-divergence and reducing the variance of confidence drops between members and non-members. Extensive experiments confirm that our method effectively mitigates these attacks, delivering privacy protection while maintaining model utility and explanation fidelity.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。