arXiv:2410.13193cs.LGcs.AI2024-10

揭示机器学习模型的'镜像陷阱',解释为何它们会误判人类不易出错的图像。

Golyadkin's Torment: Doppelgängers and Adversarial Vulnerability

  • 提出'对抗镜像'概念,用感知距离衡量输入相似性
  • 发现多数模型对镜像攻击极度脆弱,且准确率与鲁棒性难兼顾
  • 识别高准确率模型自动成为敏感型,为提升可靠性提供新路径

许多机器学习分类器声称优于人类,但仍会犯人类不会犯的错误。最典型的例子是对抗视觉同质体。本文旨在定义并研究对抗镜像(AD),包括对抗视觉同质体,并比较机器学习分类器与人类在性能和鲁棒性上的差异。我们发现,对抗镜像是相对于本文定义的感知度量而言彼此接近的输入。对抗镜像在本质上不同于常规对抗样本。绝大多数分类器对对抗镜像攻击存在漏洞,且准确率-鲁棒性权衡可能无法改善其表现。某些分类任务可能根本不存在鲁棒的对抗镜像分类器,因为底层类别本身具有模糊性。我们提供了判断分类任务是否明确定义的标准;描述了对抗镜像鲁棒分类器的结构与属性;引入并探讨了概念熵和概念模糊区域的概念,以及界定对抗镜像欺骗率的方法。我们定义了表现出超敏感行为的分类器,即其唯一错误仅为对抗镜像。提高此类超敏感分类器的对抗镜像鲁棒性等价于提升准确率。我们确定了所有高准确率分类器均为超敏感的条件。这些发现旨在显著提升机器学习系统的可靠性和安全性。

原文摘要 · Abstract (English)

Many machine learning (ML) classifiers are claimed to outperform humans, but they still make mistakes that humans do not. The most notorious examples of such mistakes are adversarial visual metamers. This paper aims to define and investigate the phenomenon of adversarial Doppelgangers (AD), which includes adversarial visual metamers, and to compare the performance and robustness of ML classifiers to human performance. We find that AD are inputs that are close to each other with respect to a perceptual metric defined in this paper. AD are qualitatively different from the usual adversarial examples. The vast majority of classifiers are vulnerable to AD and robustness-accuracy trade-offs may not improve them. Some classification problems may not admit any AD robust classifiers because the underlying classes are ambiguous. We provide criteria that can be used to determine whether a classification problem is well defined or not; describe the structure and attributes of an AD-robust classifier; introduce and explore the notions of conceptual entropy and regions of conceptual ambiguity for classifiers that are vulnerable to AD attacks, along with methods to bound the AD fooling rate of an attack. We define the notion of classifiers that exhibit hypersensitive behavior, that is, classifiers whose only mistakes are adversarial Doppelgangers. Improving the AD robustness of hyper-sensitive classifiers is equivalent to improving accuracy. We identify conditions guaranteeing that all classifiers with sufficiently high accuracy are hyper-sensitive. Our findings are aimed at significant improvements in the reliability and security of machine learning systems.

对抗攻击模型鲁棒性概念模糊感知度量

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。