用特征覆盖度选毒样本,低预算下攻击成功率超96%。
Diversity Matters: Distributional Feature Coverage Sample Selection for Data-Efficient Backdoor Attacks

- 基于预训练特征聚类,每毒槽选中心最近样本来覆盖特征分布
- 在六组实验中平均攻击成功率96.30%,最高超对比方法4.6个百分点
- 无需训练、不依赖触发器,适合资源受限的隐蔽攻击研究
后门攻击通过污染训练数据,使模型在正常输入上保持高准确率,但在特定触发输入下预测攻击者指定目标。在极低污染率下,仅有少数样本携带触发-目标关联,因此毒样本选择至关重要。现有方法通常使用单样本评分排序候选样本,易选出语义相似区域的冗余样本,且多数需任务特定的代理训练。本文提出分布特征覆盖采样(DFCS),一种无训练、触发无关的方法:将固定预训练特征聚类为每个毒槽一个区域,并从每个区域中选取距中心最近的样本。局部一阶分析表明该分配与特征覆盖和代表性质量相关。在CIFAR-10、Tiny-ImageNet和Imagenette上的BadNets与Blended攻击中,DFCS在七种选择器中平均攻击成功率最高,六组设置下均达96.30%,较最强比较方法平均提升4.60个百分点,同时保持干净准确率。结果支持分布特征覆盖作为低预算脏标签后门攻击的有效选择原则。
原文摘要 · Abstract (English)
Backdoor attacks compromise training data so that a model retains clean accuracy but predicts an attacker-chosen target on triggered inputs. At very low poisoning rates, only a few samples convey the trigger--target association, making poison-sample selection critical. Existing methods typically rank candidates using per-sample scores, which can select redundant samples from similar semantic regions, and many require task-specific surrogate training. We propose Distributional Feature Coverage Sample Selection (DFCS), a training-free, trigger-agnostic method that clusters fixed pretrained features into one region per poisoning slot and selects the centroid-nearest sample from each region. A local first-order analysis relates this allocation to feature-coverage and representative-mass terms. Across BadNets and Blended attacks on CIFAR-10, Tiny-ImageNet, and Imagenette, DFCS achieves the highest mean attack success rate among seven selectors in all six dataset--attack settings, averaging $96.30\%$ and exceeding the strongest comparator in each setting by 4.60 percentage points on average while preserving clean accuracy. These results support distributional feature coverage as an effective selection principle for low-budget dirty-label backdoor attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。