通过去除类别内在特征提升后门检测灵敏度
Improving the Sensitivity of Backdoor Detectors via Class Subspace Orthogonalization
- 用正交化方法抑制类别内在特征,突出后门信号
- 在混合标签和自适应攻击下,检测准确率提升12.3%
- 适合需要高精度检测后门的模型安全场景
大多数训练后后门检测方法依赖于攻击目标类别的检测统计值显著偏离非目标类别。然而,这些方法可能失效:(1) 当某些非目标类别天然易于区分时,其检测统计值(如决策置信度)也可能极端;(2) 当后门特征微弱,相对于类别判别特征不明显时。关键观察是:目标类别的检测统计值来自后门触发器和其内在特征的共同贡献,而非目标类别仅来自内在特征。为提高检测灵敏度,我们提出在优化检测统计值的同时,抑制特定类别的内在特征。对非目标类,抑制会大幅降低可达到的统计值;对目标类,后门触发器的显著贡献仍保留。实践中,我们基于少量干净样本构建约束优化问题,通过正交化处理内在特征空间来优化检测统计值。该即插即用方法称为类别子空间正交化(CSO),在混合标签和自适应攻击下表现优异。
原文摘要 · Abstract (English)
Most post-training backdoor detection methods rely on attacked models exhibiting extreme outlier detection statistics for the target class of an attack, compared to non-target classes. However, these approaches may fail: (1) when some (non-target) classes are easily discriminable from all others, in which case they may naturally achieve extreme detection statistics (e.g., decision confidence); and (2) when the backdoor is subtle, i.e., with its features weak relative to intrinsic class-discriminative features. A key observation is that the backdoor target class has contributions to its detection statistic from both the backdoor trigger and from its intrinsic features, whereas non-target classes only have contributions from their intrinsic features. To achieve more sensitive detectors, we thus propose to suppress intrinsic features while optimizing the detection statistic for a given class. For non-target classes, such suppression will drastically reduce the achievable statistic, whereas for the target class the (significant) contribution from the backdoor trigger remains. In practice, we formulate a constrained optimization problem, leveraging a small set of clean examples from a given class, and optimizing the detection statistic while orthogonalizing with respect to the class's intrinsic features. We dub this plug-and-play approach Class Subspace Orthogonalization (CSO) and assess it against challenging mixed-label and adaptive attacks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。