arXiv:2512.18651cs.CV2025-12

发现零样本学习模型在类别和概念层面均易受攻击,提出新攻击方法全面破坏其性能。

Adversarial Robustness in Zero-Shot Learning:An Empirical Study on Class and Concept-Level Vulnerabilities

  • 提出类偏置增强攻击(CBEA),使所有校准点下GZSL准确率归零。
  • 在类别和概念层面设计新攻击,验证现有模型普遍存在脆弱性。
  • 适用于关注模型安全性的研究人员,尤其对零样本学习应用者有警示意义。

零样本学习(ZSL)旨在让图像分类器识别训练中未出现过的类别。与传统监督分类不同,ZSL依赖于从视觉特征到预定义人类可理解类别概念的映射。尽管ZSL提升了泛化能力与可解释性,其在系统性输入扰动下的鲁棒性仍不明确。本研究通过实证分析现有ZSL方法在类别级与概念级的脆弱性。我们发现经典非目标类别攻击(clsA)虽能干扰分类结果,但在广义零样本学习(GZSL)设置下,攻击成功仅出现在最优校准点;攻击后最优校准点转移,模型在其他点仍保持较强性能,表明clsA在GZSL中为虚假成功。为此,我们提出类偏置增强攻击(CBEA),通过放大已见与未见类别概率差距,使所有校准点上的GZSL准确率降至零。此外,在概念级攻击中,引入两种新攻击模式:保留类别的概念攻击(CPconA)与不保留类别的概念攻击(NCPconA)。实验评估了近三年三种典型ZSL模型在多种架构上的表现,结果表明,现有模型不仅易受传统类别攻击,也极易被概念级攻击操控,攻击者可通过擦除或引入概念轻易篡改分类结果。研究揭示了当前方法间显著的性能差距,凸显提升ZSL模型对抗鲁棒性的迫切需求。

原文摘要 · Abstract (English)

Zero-shot Learning (ZSL) aims to enable image classifiers to recognize images from unseen classes that were not included during training. Unlike traditional supervised classification, ZSL typically relies on learning a mapping from visual features to predefined, human-understandable class concepts. While ZSL models promise to improve generalization and interpretability, their robustness under systematic input perturbations remain unclear. In this study, we present an empirical analysis about the robustness of existing ZSL methods at both classlevel and concept-level. Specifically, we successfully disrupted their class prediction by the well-known non-target class attack (clsA). However, in the Generalized Zero-shot Learning (GZSL) setting, we observe that the success of clsA is only at the original best-calibrated point. After the attack, the optimal bestcalibration point shifts, and ZSL models maintain relatively strong performance at other calibration points, indicating that clsA results in a spurious attack success in the GZSL. To address this, we propose the Class-Bias Enhanced Attack (CBEA), which completely eliminates GZSL accuracy across all calibrated points by enhancing the gap between seen and unseen class probabilities.Next, at concept-level attack, we introduce two novel attack modes: Class-Preserving Concept Attack (CPconA) and NonClass-Preserving Concept Attack (NCPconA). Our extensive experiments evaluate three typical ZSL models across various architectures from the past three years and reveal that ZSL models are vulnerable not only to the traditional class attack but also to concept-based attacks. These attacks allow malicious actors to easily manipulate class predictions by erasing or introducing concepts. Our findings highlight a significant performance gap between existing approaches, emphasizing the need for improved adversarial robustness in current ZSL models.

零样本学习对抗攻击概念攻击鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。