提出可生成流畅对抗文本的新方法,用于隐藏攻击意图。
Adversarial Decoding: Generating Readable Documents for Adversarial Objectives

- 通过对抗性解码生成连贯文本,满足多种隐蔽攻击目标
- 在RAG中毒、越狱和绕过过滤器测试中均优于现有方法
- 适用于真实场景中的间接注入攻击,如检索后诱导生成
我们设计、实现并评估了对抗性解码,一种通用的文本生成技术,能够为不同对抗性目标生成可读文档。以往方法要么产生易被识别的乱码,要么无法处理包含嵌入相似性的目标。尤其难以应对现实中的间接注入攻击,例如:(1) 能在RAG系统中响应广泛查询类别而被检索到的文档,(2) 可对后续生成产生对抗性影响。我们还发现,流畅性(低困惑度)不足以规避过滤。在多种目标下评估对抗性解码的有效性,包括RAG poisoning、jailbreaking和防御过滤绕过,结果表明其性能优于现有方法,同时生成可读的对抗文本。
原文摘要 · Abstract (English)
We design, implement, and evaluate adversarial decoding, a new, generic text generation technique that produces readable documents for different adversarial objectives. Prior methods either produce easily detectable gibberish, or cannot handle objectives that include embedding similarity. In particular, they only work for direct attacks (such as jailbreaking) and cannot produce adversarial text for realistic indirect injection, e.g., documents that (1) are retrieved in RAG systems in response to broad classes of queries, and also (2) adversarially influence subsequent generation. We also show that fluency (low perplexity) is not sufficient to evade filtering. We measure the effectiveness of adversarial decoding for different objectives, including RAG poisoning, jailbreaking, and evasion of defensive filters, and demonstrate that it outperforms existing methods while producing readable adversarial documents.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。