arXiv:2608.21803cs.CRcs.AI2026-08中稿 · the IEEE Cybersecu…

为黑箱解释模型设计零信任框架,确保解释结果真实可信。

ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models

论文配图:ExplainGuard: A Zero Trust Framework for Post-Hoc Explanation Integrity Guarantees in Blackbox XAI Models
图 1 · 摘自论文原文
  • 引入零信任架构,通过三重验证保障解释完整性。
  • 可有效防御输出混淆、分布外伪造等前沿攻击手段。
  • 适合监管合规、高风险场景下的AI解释审计需求。

随着机器学习模型在高风险环境中的广泛应用,像SHAP和LIME这样的可解释AI(XAI)方法已成为满足监管合规与建立信任的关键。然而,当前的审计范式依赖于第三方审计员“可信”的隐含假设,而近期研究揭示这一假设存在漏洞:恶意审计员可通过输出混淆或引入分布外(OOD)样本等操纵攻击,隐藏模型偏见同时维持高预测准确率,实现‘公平幻觉’解释。本文提出新型防御框架ExplainGuard,采用零信任架构(ZTA),嵌入XAI解释供应链中,确保生成解释的完整性。该框架设立策略决策点(PDP),在解释发布前执行三项验证:(1)通过行为指纹检测模型替换;(2)利用公理一致性检查排除数学上不可能的解释;(3)基于排名稳定性进行特征忠实性验证,计算开销极低。实验表明,ExplainGuard能有效中和现有最先进解释操纵攻击,并将审计过程转变为可验证操作。

原文摘要 · Abstract (English)

As machine learning (ML) models are increasingly deployed in high-stakes environments, explainable AI (XAI) methods like SHAP and LIME have become essential for regulatory compliance and trust. However, the current auditing paradigm relies on an implicit "chain of trust" where third-party auditors are assumed to be trusted. Recent research demonstrates that this assumption is flawed and adversarial auditors can manipulate XAI explanations through manipulation attacks such as output shuffling or scaffolding out-of-distribution (OOD) to conceal model biases while maintaining high prediction accuracy aiming for fairwashed explanation. In this paper, we introduce a novel defense framework, ExplainGuard, that leverages a Zero-Trust architecture (ZTA) design to be incorporated within the XAI explanation supply chain and ensures the integrity of the generated explanation. This framework would help us to replace the ambiguous default assumption of "auditor is trustworthy," with a continuous "verify-then-trust" approach. Our design architecture establishes a Policy Decision Point (PDP) that enforces three distinct pillars of verification before any explanation is released to the user: (1) asset integrity via behavioral fingerprint to detect model substitution, (2) semantic validity using axiomatic consistency checks to reject mathematically impossible explanations, and (3) feature faithfulness verification utilizing a ranking stability approach with minimal computational overhead. Finally, we evaluate how ExplainGuard can effectively neutralize state- of-the-art explanation manipulation attacks while transforming the auditing process into a verifiable operation.

可解释AI零信任模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。