用自然风格设计隐蔽文本后门,让攻击不易被人工发现。
The Ultimate Cookbook for Invisible Poison: Crafting Subtle Clean-Label Text Backdoors with Style Attributes
- 通过提取细粒度属性生成看似自然的触发词
- 人工评估显示成功率更高且更难被识别
- 揭示了自动指标与人类判断的偏差,适合安全研究者
文本分类器的后门攻击可在特定‘触发词’出现时诱使其预测预设标签。以往攻击多依赖语法错误或异常文本,易被人工标注者察觉并过滤,削弱攻击效果。本文认为成功攻击的关键是触发前后文本对人类不可区分。为此,我们首次系统开展人工评估以衡量攻击隐蔽性,并提出AttrBkd方法,包含三种生成细微但有效的触发属性的策略,例如从现有基线攻击中提取细粒度特征。人工评估表明,基于基线属性构建的AttrBkd攻击在成功率上更高,同时更隐蔽(被人类识别的实例更少),证明攻击可通过自然外观绕过检测。此外,人工标注揭示了自动化指标未能捕捉的信息,暴露出这些指标与人类判断之间的不一致。
原文摘要 · Abstract (English)
Backdoor attacks on text classifiers can cause them to predict a predefined label when a particular "trigger" is present. Prior attacks often rely on triggers that are ungrammatical or otherwise unusual, leading to conspicuous attacks. As a result, human annotators, who play a critical role in curating training data in practice, can easily detect and filter out these unnatural texts during manual inspection, reducing the risk of such attacks. We argue that a key criterion for a successful attack is for text with and without triggers to be indistinguishable to humans. However, prior work neither directly nor comprehensively evaluated attack subtlety and invisibility with human involvement. We bridge the gap by conducting thorough human evaluations to assess attack subtlety. We also propose \emph{AttrBkd}, consisting of three recipes for crafting subtle yet effective trigger attributes, such as extracting fine-grained attributes from existing baseline backdoor attacks. Our human evaluations find that AttrBkd with these baseline-derived attributes is often more effective (higher attack success rate) and more subtle (fewer instances detected by humans) than the original baseline backdoor attacks, demonstrating that backdoor attacks can bypass detection by being inconspicuous and appearing natural even upon close inspection, while still remaining effective. Our human annotation also provides information not captured by automated metrics used in prior work, and demonstrates the misalignment of these metrics with human judgment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。