针对可解释模型的隐蔽攻击,用概念触发器操控预测结果。
CAT: Concept-level backdoor ATtacks for Concept Bottleneck Models
- 在概念层嵌入触发器,训练时植入后门
- 攻击成功率高,对正常数据无影响
- 适合研究模型安全与可解释性漏洞者
尽管深度学习在多个领域产生深远影响,其内在的不透明性推动了可解释人工智能(XAI)的发展。其中,概念瓶颈模型(CBM)通过利用高层语义信息,成为提升可解释性的关键方法。然而,与其它机器学习模型一样,CBM也面临安全威胁,尤其是后门攻击,可能在暗中操控模型行为。由于学界尚未研究CBM的概念级后门攻击,我们提出CAT(概念级后门攻击)方法,利用CBM中的概念表示在训练阶段嵌入触发器,实现在推理时对模型预测的可控操纵。增强版攻击模式CAT+引入相关性函数,系统选择最有效且隐蔽的概念触发器,优化攻击效果。我们的综合评估框架衡量攻击成功率与隐蔽性,结果显示CAT和CAT+在干净数据上保持高性能,同时在被污染数据上实现显著的目标性影响。本工作揭示了CBM潜在的安全风险,并为未来安全评估提供了可靠测试方法。
原文摘要 · Abstract (English)
Despite the transformative impact of deep learning across multiple domains, the inherent opacity of these models has driven the development of Explainable Artificial Intelligence (XAI). Among these efforts, Concept Bottleneck Models (CBMs) have emerged as a key approach to improve interpretability by leveraging high-level semantic information. However, CBMs, like other machine learning models, are susceptible to security threats, particularly backdoor attacks, which can covertly manipulate model behaviors. Understanding that the community has not yet studied the concept level backdoor attack of CBM, because of "Better the devil you know than the devil you don't know.", we introduce CAT (Concept-level Backdoor ATtacks), a methodology that leverages the conceptual representations within CBMs to embed triggers during training, enabling controlled manipulation of model predictions at inference time. An enhanced attack pattern, CAT+, incorporates a correlation function to systematically select the most effective and stealthy concept triggers, thereby optimizing the attack's impact. Our comprehensive evaluation framework assesses both the attack success rate and stealthiness, demonstrating that CAT and CAT+ maintain high performance on clean data while achieving significant targeted effects on backdoored datasets. This work underscores the potential security risks associated with CBMs and provides a robust testing methodology for future security assessments.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。