arXiv:2512.12827cs.LGcs.CR2025-12

通过梯度内在维度差异,高效识别对抗样本。

GradID: Adversarial Detection via Intrinsic Dimensionality of Gradients

  • 利用梯度内在维度差异检测对抗样本
  • 在CIFAR-10上对多种攻击检测率超92%
  • 适用于批量和单样本场景,适合高可靠性应用

尽管深度神经网络表现优异,但其对微小且通常不可察觉的对抗扰动极为敏感,可能导致预测结果大幅偏离。鉴于医疗诊断与自动驾驶等应用对可靠性的严苛要求,有效检测对抗攻击至关重要。本文研究模型输入损失景观的几何特性,分析梯度参数的内在维度(ID),即描述数据点所在流形所需最少坐标数。我们发现自然数据与对抗数据在ID上存在显著且一致的差异,据此提出新的检测方法。在两类典型场景中验证:一是批量场景下识别恶意数据组,在MNIST与SVHN上表现优异;二是关键的单样本场景中,在CIFAR-10与MS COCO等挑战性基准上达到新SOTA。该方法在多种攻击(如CW、AutoAttack)下显著优于现有方法,于CIFAR-10上检测率持续高于92%。结果表明,内在维度是一种强大且稳健的对抗样本指纹,适用于多种数据集与攻击策略。

原文摘要 · Abstract (English)

Despite their remarkable performance, deep neural networks exhibit a critical vulnerability: small, often imperceptible, adversarial perturbations can lead to drastically altered model predictions. Given the stringent reliability demands of applications such as medical diagnosis and autonomous driving, robust detection of such adversarial attacks is paramount. In this paper, we investigate the geometric properties of a model's input loss landscape. We analyze the Intrinsic Dimensionality (ID) of the model's gradient parameters, which quantifies the minimal number of coordinates required to describe the data points on their underlying manifold. We reveal a distinct and consistent difference in the ID for natural and adversarial data, which forms the basis of our proposed detection method. We validate our approach across two distinct operational scenarios. First, in a batch-wise context for identifying malicious data groups, our method demonstrates high efficacy on datasets like MNIST and SVHN. Second, in the critical individual-sample setting, we establish new state-of-the-art results on challenging benchmarks such as CIFAR-10 and MS COCO. Our detector significantly surpasses existing methods against a wide array of attacks, including CW and AutoAttack, achieving detection rates consistently above 92\% on CIFAR-10. The results underscore the robustness of our geometric approach, highlighting that intrinsic dimensionality is a powerful fingerprint for adversarial detection across diverse datasets and attack strategies.

对抗检测内在维度梯度分析鲁棒性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。