arXiv:2602.16494cs.CV2026-02

提出统一评估框架,揭示检测模型对抗攻击的迁移性缺陷及最优防御策略。

Benchmarking Adversarial Robustness and Adversarial Training Strategies for Object Detection

  • 构建统一基准,分离定位与分类错误,多维度评估攻击成本。
  • 攻击对视觉变换器迁移效果差,卷积模型攻击难泛化到新架构。
  • 混合高扰动攻击(空间+语义)训练最有效,优于单一攻击训练。

目标检测模型是自动驾驶和感知机器人等自动化系统的核心组件,但其对对抗攻击的敏感性带来严重安全风险。当前防御研究进展滞后于分类任务,主要受限于缺乏标准化评估体系——现有工作使用不同数据集、不一致的效率指标和多样的扰动成本衡量方式,难以全面比较攻击或防御方法。本文通过三个核心问题展开:(1) 如何建立公平的基准以公正比较攻击?(2) 现代攻击在不同架构间(特别是从卷积神经网络到视觉变换器)的迁移能力如何?(3) 哪种对抗训练策略最有效?为此,我们提出一个聚焦数字域、非贴片类攻击的统一基准框架,引入特定指标分离定位与分类误差,并采用多种感知度量评估攻击成本。基于该框架,我们在先进攻击和广泛检测器上进行大量实验。结果表明:第一,现代对抗攻击在卷积模型与视觉变换器之间具有显著迁移性缺失;第二,最鲁棒的对抗训练策略是使用包含多种高扰动攻击(如空间与语义目标)混合的数据集,其性能优于仅使用单一攻击的训练方式。

原文摘要 · Abstract (English)

Object detection models are critical components of automated systems, such as autonomous vehicles and perception-based robots, but their sensitivity to adversarial attacks poses a serious security risk. Progress in defending these models lags behind classification, hindered by a lack of standardized evaluation. It is nearly impossible to thoroughly compare attack or defense methods, as existing work uses different datasets, inconsistent efficiency metrics, and varied measures of perturbation cost. This paper addresses this gap by investigating three key questions: (1) How can we create a fair benchmark to impartially compare attacks? (2) How well do modern attacks transfer across different architectures, especially from Convolutional Neural Networks to Vision Transformers? (3) What is the most effective adversarial training strategy for robust defense? To answer these, we first propose a unified benchmark framework focused on digital, non-patch-based attacks. This framework introduces specific metrics to disentangle localization and classification errors and evaluates attack cost using multiple perceptual metrics. Using this benchmark, we conduct extensive experiments on state-of-the-art attacks and a wide range of detectors. Our findings reveal two major conclusions: first, modern adversarial attacks against object detection models show a significant lack of transferability to transformer-based architectures. Second, we demonstrate that the most robust adversarial training strategy leverages a dataset composed of a mix of high-perturbation attacks with different objectives (e.g., spatial and semantic), which outperforms training on any single attack.

目标检测对抗攻击鲁棒性视觉变换器

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。