arXiv:2605.25304cs.LGcs.CR2026-05中稿 · CVPR

可解释模型的语义层竟成攻击突破口,新方法显著提升防御能力。

When Interpretability Becomes a Liability: Adversarial Attacks on CBM Concept Layers

论文配图:When Interpretability Becomes a Liability: Adversarial Attacks on CBM Concept Layers
图 1 · 摘自论文原文
  • 通过概念空间扰动,发现输入微小变化即可导致分类崩溃。
  • 标准CBM在CUB数据集上对概念攻击极度脆弱,最小扰动仅0.46。
  • 提出SPECTRA防御机制,使攻击所需扰动增大超9000倍,精度损失<2.2%。

概念瓶颈模型(CBMs)是可解释机器学习的核心方法,通过显式概念激活提供人类可理解的中间表示。然而,这种可解释性引入了此前未被探索的攻击面:概念瓶颈层本身。本文系统研究了CBMs在概念层面的对抗脆弱性,发现针对输入像素的微小、有针对性的扰动可操纵语义表征,引发灾难性误分类。我们构建了严格的理论框架,量化概念空间鲁棒性,提出新指标揭示架构脆弱性。在CUB-200-2011数据集上的广泛分析表明,标准CBMs对概念级操控极为敏感。为此,我们提出SPECTRA(基于语义扰动的概念训练以增强抗攻击能力),一种原则性的稳定性正则化防御方法。SPECTRA有效硬化语义表示空间,使成功攻击所需的最小扰动范数从0.46提升至超过4,200,使目标概念操纵在计算上不可行。同时,保持基线分类精度在2.2%以内。本工作确立概念级攻击为一类根本性威胁模型,开启可解释学习与对抗鲁棒性交叉研究的新前沿。

原文摘要 · Abstract (English)

Concept Bottleneck Models (CBMs) have emerged as a cornerstone approach for interpretable machine learning, providing human-understandable intermediate representations through explicit concept activations. However, this interpretability fundamentally introduces a critical, previously unexplored attack surface: the concept bottleneck layer itself. We present a comprehensive, systematic study of concept-level adversarial vulnerabilities in CBMs, revealing that targeted, minimal perturbations operating on input pixels can induce catastrophic misclassification by manipulating semantic representations. We develop a rigorous theoretical framework to quantify concept-space robustness, establishing novel metrics that expose the vulnerability landscape of these architectures. Our extensive analysis on the CUB-200-2011 dataset demonstrates that standard CBMs exhibit severe susceptibility to concept-level manipulation. To address this critical weakness, we introduce SPECTRA (Semantic Perturbation-based Concept Training for Robustness against Attacks), a principled stability regularization defense. SPECTRA effectively hardens the semantic representation space, increasing the minimal perturbation norm required for a successful attack from 0.46 to over 4,200, rendering targeted concept manipulation computationally prohibitive. Furthermore, SPECTRA preserves baseline classification accuracy to within 2.2%. By establishing concept-level attacks as a fundamentally distinct threat model, this work opens a new research frontier at the intersection of interpretable machine learning and adversarial robustness.

可解释AI对抗攻击语义鲁棒性概念模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。