攻击视觉基础模型的通用对抗样本,破坏其特征表示。
Task-Agnostic Attacks Against Vision Foundation Models
- 通过最大化干扰特征表示,生成不依赖具体任务的对抗样本。
- 攻击在多个下游任务上均有效,且跨模型迁移性强。
- 适合关注视觉模型安全性的研究人员和开发者。
机器学习安全研究主要集中在下游任务特定的攻击,即针对特定任务优化损失函数以生成对抗样本。与此同时,机器学习从业者普遍采用公开预训练的视觉基础模型,这些模型共享相同的骨干架构,广泛应用于分类、分割、深度估计、检索、问答等多种任务。然而,对这类基础模型的攻击及其对多下游任务的影响仍远未得到充分研究。本文提出一种通用框架,通过最大程度地破坏视觉基础模型的特征表示,生成任务无关的对抗样本。我们通过测量该攻击在多个下游任务上的影响及其在不同模型间的可迁移性,系统评估了主流视觉基础模型的特征表示安全性。
原文摘要 · Abstract (English)
The study of security in machine learning mainly focuses on downstream task-specific attacks, where the adversarial example is obtained by optimizing a loss function specific to the downstream task. At the same time, it has become standard practice for machine learning practitioners to adopt publicly available pre-trained vision foundation models, effectively sharing a common backbone architecture across a multitude of applications such as classification, segmentation, depth estimation, retrieval, question-answering and more. The study of attacks on such foundation models and their impact to multiple downstream tasks remains vastly unexplored. This work proposes a general framework that forges task-agnostic adversarial examples by maximally disrupting the feature representation obtained with foundation models. We extensively evaluate the security of the feature representations obtained by popular vision foundation models by measuring the impact of this attack on multiple downstream tasks and its transferability between models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。