arXiv:2502.19502cs.LG2025-02被引 1

提出可解释但不透明的模型保护方法,让人类能懂决策逻辑,黑客却难复制。

Models That Are Interpretable But Not Transparent

  • 基于最大集合覆盖的优化框架,生成完全忠实的解释。
  • 在保证解释完全真实的同时,最小化决策边界的泄露信息。
  • 适合需要可解释性又怕被逆向攻击的高风险场景使用。

在高风险应用中,模型的可信解释至关重要。固有可解释模型因其天然揭示决策逻辑而适合此类场景,但模型设计者常需保密以维护其价值。这引发矛盾:既要模型可解释(便于人类理解与验证预测),又要不透明(防止攻击者复刻决策边界)。然而,完全忠实的解释会暴露每个查询点附近子空间的完整逻辑,加剧泄露风险。本文提出FaithfulDefense方法,能在确保解释完全忠实的前提下,最大限度减少对决策边界的揭示。该方法基于最大集合覆盖的优化问题,并利用子模性设计多种求解形式。

原文摘要 · Abstract (English)

Faithful explanations are essential for machine learning models in high-stakes applications. Inherently interpretable models are well-suited for these applications because they naturally provide faithful explanations by revealing their decision logic. However, model designers often need to keep these models proprietary to maintain their value. This creates a tension: we need models that are interpretable--allowing human decision-makers to understand and justify predictions, but not transparent, so that the model's decision boundary is not easily replicated by attackers. Shielding the model's decision boundary is particularly challenging alongside the requirement of completely faithful explanations, since such explanations reveal the true logic of the model for an entire subspace around each query point. This work provides an approach, FaithfulDefense, that creates model explanations for logical models that are completely faithful, yet reveal as little as possible about the decision boundary. FaithfulDefense is based on a maximum set cover formulation, and we provide multiple formulations for it, taking advantage of submodularity.

可解释性模型安全隐私保护

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。