arXiv:2503.06529cs.CRcs.AI2025-03被引 4

提出可任意控制目标的检测模型后门攻击,实现对象消失、伪造或误标。

AnywhereDoor: Multi-Target Backdoor Attacks on Object Detection

  • 通过解耦目标实现多目标灵活控制,突破单一目标限制。
  • 攻击成功率提升26%,在多种检测器上表现稳定。
  • 适合研究模型安全与对抗防御的人员阅读。

随着目标检测广泛应用于诸多关键场景,理解其漏洞至关重要。后门攻击通过在受害模型中植入隐藏触发器,在推理时引发恶意行为,构成严重威胁。然而现有研究仅限于单目标攻击,攻击者需预先定义固定目标,无法实现推理时动态调整。由于目标检测输出空间庞大(包含存在性预测、边界框估计和分类),实现灵活、推理时可控的攻击尚未被探索。本文提出 AnywhereDoor,一种面向目标检测的多目标后门攻击。一旦植入,攻击者可使对象消失、虚构新对象或错误标注,覆盖所有类别或特定类别,实现前所未有的控制能力。该灵活性由三项创新实现:(i) 目标解耦以扩展支持目标数量;(ii) 触发器马赛克化增强对区域检测器的鲁棒性;(iii) 战略性分批处理缓解对象级数据不平衡问题。大量实验表明,AnywhereDoor显著提升攻击控制力,相比现有方法适配版本,攻击成功率提升26%。

原文摘要 · Abstract (English)

As object detection becomes integral to many safety-critical applications, understanding its vulnerabilities is essential. Backdoor attacks, in particular, pose a serious threat by implanting hidden triggers in victim models, which adversaries can later exploit to induce malicious behaviors during inference. However, current understanding is limited to single-target attacks, where adversaries must define a fixed malicious behavior (target) before training, making inference-time adaptability impossible. Given the large output space of object detection (including object existence prediction, bounding box estimation, and classification), the feasibility of flexible, inference-time model control remains unexplored. This paper introduces AnywhereDoor, a multi-target backdoor attack for object detection. Once implanted, AnywhereDoor allows adversaries to make objects disappear, fabricate new ones, or mislabel them, either across all object classes or specific ones, offering an unprecedented degree of control. This flexibility is enabled by three key innovations: (i) objective disentanglement to scale the number of supported targets; (ii) trigger mosaicking to ensure robustness even against region-based detectors; and (iii) strategic batching to address object-level data imbalances that hinder manipulation. Extensive experiments demonstrate that AnywhereDoor grants attackers a high degree of control, improving attack success rates by 26% compared to adaptations of existing methods for such flexible control.

后门攻击目标检测模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。