让后门样本藏进数据稀疏区,躲过模型修复
Density-aware Sample-specific Attack

- 通过密度感知优化,将触发样本放入数据低密度区
- 攻击成功率超99%,防御后仍比基线高50-85个百分点
- 对剪枝防御完全免疫,适合研究新型防御机制者
尽管后门攻击取得进展,现有方法仍易被训练后防御(如微调或剪枝)消除。本文重新审视后门攻击目标,基于贝叶斯最优模型推导出样本特定触发器的最优构造准则:当受污染样本被引导至干净数据分布的低密度区域时,攻击成功率与纯净准确率可同时优化。该分布条件一次性控制了污染分布的所有阶矩,而非仅依赖输入空间的少量统计量。我们提出双层优化框架,通过条件时间得分匹配估计密度比,并优化混合模型目标,将触发样本置于这些稀疏区域。在MNIST、CIFAR-10、GTSRB和TinyImageNet上的大量实验表明,该方法在无防御时攻击成功率超过99%,在微调防御下仍保持比最强基线高50–85个百分点的后防御攻击成功率。面对神经元剪枝防御,该方法表现出完全免疫性,所有剪枝阈值下均无神经元被识别为可移除。结果揭示了当前防御范式的根本缺陷,强调需发展超越干净数据支撑集的防御策略。
原文摘要 · Abstract (English)
Despite recent progress in backdoor attacks, existing methods remain susceptible to post-training defenses that erase the backdoor through fine-tuning or pruning. We revisit the core objectives of backdoor attacks and derive principled criteria characterizing optimal sample-specific trigger construction under a Bayes-optimal model of the victim's training. Our analysis reveals that both attack success and clean-accuracy preservation are simultaneously optimized when triggered samples are steered into low-density regions of the clean data distribution, a distributional condition that controls all moments of the poisoned distribution at once rather than a handful of input-space summary statistics. We introduce a bilevel optimization framework that estimates density ratios via conditional time-score matching and optimizes a mixture-model objective to place triggered samples in these sparse regions. Extensive evaluations on MNIST, CIFAR-10, GTSRB, and TinyImageNet demonstrate that our method achieves above 99\% attack success rate before defense and retains 50--85 percentage points higher post-defense ASR than the strongest baselines under fine-tuning defenses. Against neuron-pruning defenses, the method exhibits complete immunity, with zero neurons identified for removal across all pruning thresholds. These results expose a fundamental gap in current defense paradigms and underscore the need for defenses that operate beyond the support of the clean distribution.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。