用自然语义藏毒,让语言模型悄悄中招
SteganoBackdoor: Stealthy and Data-Efficient Backdoor Attacks on Language Models
- 把后门代码藏进流畅句子,不改变原意
- 仅用少量数据就能让模型精准触发攻击
- 适合研究安全防御的学者和工程师
现代语言模型仍易受投毒数据引发的后门攻击,即训练时在含触发词的输入与目标输出间建立关联,导致推理时只要出现触发词就触发恶意行为。现有研究多关注通过风格化痕迹或词元级扰动实现隐蔽攻击,但忽视了更贴近实际的威胁:与自然语义概念绑定的后门。我们提出SteganoBackdoor,一种基于优化的框架,生成称为SteganoPoisons的隐写投毒样本——将后门载荷分散嵌入流畅句子中,且与推理时的语义触发器无表征重叠。在多种模型架构上,该方法在有限投毒预算下仍保持高攻击成功率,并在保守的数据过滤条件下依然有效,揭示了当前数据清洗防御机制的盲点。
原文摘要 · Abstract (English)
Modern language models remain vulnerable to backdoor attacks via poisoned data, where training inputs containing a trigger are paired with a target output, causing the model to reproduce that behavior whenever the trigger appears at inference time. Recent work has emphasized stealthy attacks that stress-test data-curation defenses using stylized artifacts or token-level perturbations as triggers, but this focus leaves a more practically relevant threat model underexplored: backdoors tied to naturally occurring semantic concepts. We introduce SteganoBackdoor, an optimization-based framework that constructs SteganoPoisons, steganographic poisoned training examples in which a backdoor payload is distributed across a fluent sentence while exhibiting no representational overlap with the inference-time semantic trigger. Across diverse model architectures, SteganoBackdoor achieves high attack success under constrained poisoning budgets and remains effective under conservative data-level filtering, highlighting a blind spot in existing data-curation defenses.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。