揭示CAV方法的理论缺陷并提出针对性攻击
Concept activation vectors: a unifying view and adversarial attacks
- 从概率视角重新理解CAV,将其视为随机向量
- 发现CAV结果高度依赖非概念数据分布
- 提出简单有效的对抗攻击,警示解释可靠性
概念激活向量(CAV)是可解释人工智能中的工具,用于理解人类可理解的概念如何编码在模型的隐空间中。它通过概念类和非概念样本的隐藏层激活计算得出。本文从概率视角出发,发现非概念输入的分布会诱导出CAV的分布,使其成为隐空间中的随机向量,从而可推导出不同类型的CAV的均值与协方差,形成统一的理论框架。这一视角也揭示了潜在漏洞:CAV可能强烈依赖于人为选择的非概念分布,该因素在以往研究中被忽视。文中以一个简单但有效的对抗攻击为例,验证了该脆弱性,强调需对CAV的稳健性进行系统性研究。
原文摘要 · Abstract (English)
Concept Activation Vectors (CAVs) are a tool from explainable AI, offering a promising approach for understanding how human-understandable concepts are encoded in a model's latent spaces. They are computed from hidden-layer activations of inputs belonging either to a concept class or to non-concept examples. Adopting a probabilistic perspective, the distribution of the (non-)concept inputs induces a distribution over the CAV, making it a random vector in the latent space. This enables us to derive mean and covariance for different types of CAVs, leading to a unified theoretical view. This probabilistic perspective also reveals a potential vulnerability: CAVs can strongly depend on the rather arbitrary non-concept distribution, a factor largely overlooked in prior work. We illustrate this with a simple yet effective adversarial attack, underscoring the need for a more systematic study.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。