arXiv:2411.16512cs.CRcs.CV2024-11被引 3

提出防御概念瓶颈模型中隐蔽后门攻击的新方法,提升高风险场景下的AI可信度。

Guarding the Gate: ConceptGuard Battles Concept-Level Backdoors in Concept Bottleneck Models

  • 通过文本距离聚类与分类器投票机制识别并隔离概念级后门触发器。
  • 在多种数据集上实现90%以上准确率的同时有效抵御后门攻击。
  • 专为可解释模型设计,适合医疗等高可靠性要求的AI应用。

随着深度学习模型复杂度上升,其透明性与可问责性引发关注,尤其在医疗诊断等高风险场景中,黑箱模型会削弱信任。可解释人工智能(XAI)旨在提供清晰可解释的模型。其中,概念瓶颈模型(CBM)通过使用高层次语义概念增强透明性。然而,CBM易受概念级后门攻击,此类攻击将隐藏触发器注入概念中,导致难以察觉的异常行为。为此,本文提出ConceptGuard,一种专为防御CBM中概念级后门攻击而设计的新型防御框架。ConceptGuard采用多阶段策略,包括基于文本距离的概念聚类及在不同概念子组上训练的分类器投票机制,以隔离并缓解潜在触发器。主要贡献有三:(i) ConceptGuard是首个针对CBM中概念级后门攻击的专门防御机制;(ii) 提供理论保障,证明其在一定触发器规模阈值内能有效防御;(iii) 实验表明,ConceptGuard在保持CBM高精度与可解释性的前提下显著提升安全性。通过全面实验与理论证明,证实ConceptGuard大幅增强了CBM的安全性与可信度,推动其在关键领域的安全部署。

原文摘要 · Abstract (English)

The increasing complexity of AI models, especially in deep learning, has raised concerns about transparency and accountability, particularly in high-stakes applications like medical diagnostics, where opaque models can undermine trust. Explainable Artificial Intelligence (XAI) aims to address these issues by providing clear, interpretable models. Among XAI techniques, Concept Bottleneck Models (CBMs) enhance transparency by using high-level semantic concepts. However, CBMs are vulnerable to concept-level backdoor attacks, which inject hidden triggers into these concepts, leading to undetectable anomalous behavior. To address this critical security gap, we introduce ConceptGuard, a novel defense framework specifically designed to protect CBMs from concept-level backdoor attacks. ConceptGuard employs a multi-stage approach, including concept clustering based on text distance measurements and a voting mechanism among classifiers trained on different concept subgroups, to isolate and mitigate potential triggers. Our contributions are threefold: (i) we present ConceptGuard as the first defense mechanism tailored for concept-level backdoor attacks in CBMs; (ii) we provide theoretical guarantees that ConceptGuard can effectively defend against such attacks within a certain trigger size threshold, ensuring robustness; and (iii) we demonstrate that ConceptGuard maintains the high performance and interpretability of CBMs, crucial for trustworthiness. Through comprehensive experiments and theoretical proofs, we show that ConceptGuard significantly enhances the security and trustworthiness of CBMs, paving the way for their secure deployment in critical applications.

可解释AI后门攻击概念瓶颈安全防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。