arXiv:2511.12722cs.LG2025-11AAAI被引 1

评估线性分类器对定向数据投毒的鲁棒性,提出高效上下界计算方法。

On Robustness of Linear Classifiers to Targeted Data Poisoning

  • 基于标签扰动的威胁模型,证明鲁棒性计算为NP完全问题。
  • 提出可高效求解上下界的算法,在多个公开数据集上验证有效。
  • 适用于现有方法失效的场景,适合安全评估与防御研究者使用。

数据投毒是训练阶段的攻击,会破坏模型的可信性。在定向数据投毒中,攻击者篡改训练数据以改变特定测试样本的分类结果。由于训练数据通常规模庞大,人工检测投毒困难。本文聚焦于自动度量数据集对这类攻击的鲁棒性。我们考虑一种威胁模型:攻击者仅能修改训练数据的标签,且仅知目标模型的假设空间。在此设定下,我们证明即使假设为线性分类器,鲁棒性计算仍是NP完全问题。为此,我们提出一种技术,可高效求得鲁棒性的下界和上界。其实现能对多个公开数据集快速计算这些边界。实验表明,当投毒程度超过识别出的鲁棒性边界时,测试点分类显著受影响。此外,我们的方法在许多当前先进方法无法处理的情况下仍可适用。

原文摘要 · Abstract (English)

Data poisoning is a training-time attack that undermines the trustworthiness of learned models. In a targeted data poisoning attack, an adversary manipulates the training dataset to alter the classification of a targeted test point. Given the typically large size of training dataset, manual detection of poisoning is difficult. An alternative is to automatically measure a dataset's robustness against such an attack, which is the focus of this paper. We consider a threat model wherein an adversary can only perturb the labels of the training dataset, with knowledge limited to the hypothesis space of the victim's model. In this setting, we prove that finding the robustness is an NP-Complete problem, even when hypotheses are linear classifiers. To overcome this, we present a technique that finds lower and upper bounds of robustness. Our implementation of the technique computes these bounds efficiently in practice for many publicly available datasets. We experimentally demonstrate the effectiveness of our approach. Specifically, a poisoning exceeding the identified robustness bounds significantly impacts test point classification. We are also able to compute these bounds in many more cases where state-of-the-art techniques fail.

数据安全投毒攻击鲁棒性评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。