arXiv:2605.01298cs.CRcs.CV2026-05

提出无需数据和训练的棋盘式后门攻击,效果优于现有方法。

Checkerboard: Closed-Form and Data-Independent Trigger Design for Clean-Label Backdoor Attacks

论文配图:Checkerboard: Closed-Form and Data-Independent Trigger Design for Clean-Label Backdoor Attacks
图 1 · 摘自论文原文
  • 基于输入空间可分性设计棋盘图案触发器,闭式求解无需优化。
  • 在CIFAR-10上仅20样本中毒即达95.72%攻击成功率。
  • 对主流防御有效,适合研究模型安全与对抗攻击者。

后门攻击通过污染少量训练数据,使模型在干净输入下正常运行,但在触发输入下被引导至攻击者指定类别。清洁标签攻击尤其难以检测,因其污染样本保持原始语义标签。然而现有方法常需代理模型训练、辅助数据或迭代优化。本文提出 extit{Checkerboard},一种闭式、数据无关的清洁标签后门攻击。通过输入空间的Fisher可分性目标建模,并在自然图像四邻域平滑先验下,得到无需数据访问、模型训练或优化的像素级棋盘触发器。在四个基准数据集上,Checkerboard优于评估的范数有界清洁标签攻击,在低全局中毒率下达到最先进性能。例如,在CIFAR-10上,触发扰动为 $10/255$,毒化20个样本即可实现95.72%攻击成功率(ASR);在IN-100上,仅0.4%全局中毒率即可获得超83%的ASR,且不降低干净准确率。该攻击对当前主流后门防御仍有效,并可通过简单修改抵抗自适应防御。

原文摘要 · Abstract (English)

Backdoor attacks threaten the deep-learning supply chain by poisoning a small fraction of the training data so that a model behaves normally on clean inputs but maps triggered inputs to an attacker-chosen class. Clean-label backdoor attacks are especially difficult to audit because poisoned examples preserve their semantic labels. Yet existing clean-label attacks often require surrogate model training, auxiliary data access, or iterative optimization. In this paper, we present \emph{Checkerboard}, a clean-label backdoor attack with closed-form, data-independent trigger design. We formulate trigger design through an input-space Fisher-separability objective and, under a ridge four-neighbor local-smoothness prior for natural images, obtain the pixel-wise checkerboard as a closed-form maximizer of the resulting design proxy without data access, model training, or optimization. Across four benchmark datasets, Checkerboard outperforms the evaluated norm-bounded clean-label attacks and achieves state-of-the-art performance under low global poisoning rates. For example, on CIFAR-10, under a trigger perturbation of $10/255$, poisoning 20 training samples achieves $95.72\%$ Attack Success Rate (ASR). On IN-100, a global poisoning rate of only $0.4\%$ yields over $83\%$ ASR without degrading clean accuracy. The proposed attack also remains effective against state-of-the-art backdoor defenses and shows resistance to adaptive defenses under simple modification.

后门攻击清洁标签触发器设计模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。