arXiv:2602.23117cs.CVcs.AI2026-02综述被引 6

系统梳理对抗样本迁移性,建立评估基准。

Devling into Adversarial Transferability on Image Classification: Review, Benchmark, and Evaluation

  • 按攻击策略将迁移性攻击分为六类
  • 提出标准化评估框架避免对比偏差
  • 适合安全研究者与模型防御开发者

对抗迁移性指在替代模型上生成的对抗样本能有效欺骗未暴露的受害者模型。该特性使攻击无需直接访问目标模型,引发重大安全风险,近年受到广泛关注。本文指出当前缺乏统一评估框架与标准,导致对现有方法的评价可能存在偏倚。为此,我们系统回顾了数百篇相关工作,将各类基于迁移的攻击归纳为六种类型;提出一个全面的评估框架,作为基准用于公正衡量不同方法;同时总结提升迁移性的常见策略,并揭示可能导致不公平比较的普遍问题。最后,简要综述了图像分类之外的迁移性攻击研究。

原文摘要 · Abstract (English)

Adversarial transferability refers to the capacity of adversarial examples generated on the surrogate model to deceive alternate, unexposed victim models. This property eliminates the need for direct access to the victim model during an attack, thereby raising considerable security concerns in practical applications and attracting substantial research attention recently. In this work, we discern a lack of a standardized framework and criteria for evaluating transfer-based attacks, leading to potentially biased assessments of existing approaches. To rectify this gap, we have conducted an exhaustive review of hundreds of related works, organizing various transfer-based attacks into six distinct categories. Subsequently, we propose a comprehensive framework designed to serve as a benchmark for evaluating these attacks. In addition, we delineate common strategies that enhance adversarial transferability and highlight prevalent issues that could lead to unfair comparisons. Finally, we provide a brief review of transfer-based attacks beyond image classification.

对抗样本迁移性安全评估综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。