arXiv:2412.05010cs.LGcs.AI2024-12被引 3

提出新型后门攻击,专攻异常检测模型的开放集性能

Backdooring Outlier Detection Methods: A Novel Attack Approach

  • 设计两类触发器,使正常样本被误判为异常,反之亦然
  • 在多个真实数据集上显著降低异常检测模型性能,防御后仍有效
  • 针对自动驾驶、医疗影像等关键场景中的模型可靠性威胁

现有后门攻击主要针对分类器的闭集性能,但忽略了对开放集性能——即异常检测——的威胁。可靠的异常检测对自动驾驶、医学影像分析等关键应用至关重要。本文指出,现有攻击无法有效影响开放集性能,因其仅干扰闭集内部决策边界。为此,我们提出BATOD:一种面向异常检测任务的新式后门攻击。通过设计两类触发器,分别将正常样本误导为异常样本,或将异常样本伪装为正常样本。我们在多个真实数据集上评估了BATOD,结果表明其在攻击前后均显著削弱分类器的开放集性能,优于已有攻击方法。

原文摘要 · Abstract (English)

There have been several efforts in backdoor attacks, but these have primarily focused on the closed-set performance of classifiers (i.e., classification). This has left a gap in addressing the threat to classifiers' open-set performance, referred to as outlier detection in the literature. Reliable outlier detection is crucial for deploying classifiers in critical real-world applications such as autonomous driving and medical image analysis. First, we show that existing backdoor attacks fall short in affecting the open-set performance of classifiers, as they have been specifically designed to confuse intra-closed-set decision boundaries. In contrast, an effective backdoor attack for outlier detection needs to confuse the decision boundary between the closed and open sets. Motivated by this, in this study, we propose BATOD, a novel Backdoor Attack targeting the Outlier Detection task. Specifically, we design two categories of triggers to shift inlier samples to outliers and vice versa. We evaluate BATOD using various real-world datasets and demonstrate its superior ability to degrade the open-set performance of classifiers compared to previous attacks, both before and after applying defenses.

后门攻击异常检测开放集学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。