现有后门检测方法在真实场景中可靠性差,因训练强度影响检测效果。
Rethinking Backdoor Detection Evaluation for Language Models
- 通过调节中毒数据训练强度,测试检测器鲁棒性
- 高强度或低强度训练的后门更难被发现
- 适合关注模型安全评估的研究者和实践者
后门攻击会使语言模型在接收到特定触发词时表现出恶意行为,对依赖公开模型的从业者构成重大安全威胁。为此,后门检测方法旨在识别发布模型是否包含后门。尽管现有检测方法在标准基准上表现准确,但其在真实场景中的鲁棒性尚不明确。本文通过操纵后门植入过程中的不同因素,检验检测器的鲁棒性。结果表明,基于触发词反演或元分类器的方法对模型在污染数据上的训练强度极为敏感:以更激进或更保守的方式训练的后门,比默认情况下的后门显著更难检测。研究揭示了现有检测方法的脆弱性以及当前基准构建的局限性。
原文摘要 · Abstract (English)
Backdoor attacks, in which a model behaves maliciously when given an attacker-specified trigger, pose a major security risk for practitioners who depend on publicly released language models. As a countermeasure, backdoor detection methods aim to detect whether a released model contains a backdoor. While existing backdoor detection methods have high accuracy in detecting backdoored models on standard benchmarks, it is unclear whether they can robustly identify backdoors in the wild. In this paper, we examine the robustness of backdoor detectors by manipulating different factors during backdoor planting. We find that the success of existing methods based on trigger inversion or meta classifiers highly depends on how intensely the model is trained on poisoned data. Specifically, backdoors planted with more aggressive or more conservative training are significantly more difficult to detect than the default ones. Our results highlight a lack of robustness of existing backdoor detectors and the limitations in current benchmark construction.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。