arXiv:2606.18530cs.CRcs.CL2026-06被引 1

评估五种提示防御在隐蔽域攻击下的效果,发现重述内容最有效。

Evaluating Prompting-Based Defenses Against Domain-Camouflaged Injection Attacks

  • 通过重述检索内容来削弱隐蔽攻击
  • 重述使攻击成功率降低55%-84%,优于现有防护配置
  • 效果因模型而异,金融领域风险最高

域隐蔽注入攻击利用领域适配词汇嵌入恶意指令,逃避依赖语法标记的标准检测。当检测失效时,需明确何种防御架构可降低攻击成功率。本文在三个模型家族(Claude Haiku、Llama 3.1 8B、Gemini 2.0 Flash)和三个部署领域(金融、法律、通用)上,对五种提示防御(聚焦、重述、提示夹心及两种组合)进行了3,510次测试。结果表明,处理前重述检索内容是最稳定有效的防御,使攻击成功率下降55%-84%(依模型而定),且在所有测试模型上均优于Llama Guard 4配置。防御效果高度依赖模型:聚焦法在Claude Haiku上使攻击率减半,但在Llama 3.1 8B上无效。金融领域基础攻击成功率高达26%-33%,无一种提示防御能完全消除弱模型上的威胁。本研究首次系统评估提示防御对隐蔽类注入攻击的应对能力,为实践提供基准建议。所有测试基于合成专业文档,其结论是否适用于真实企业文档仍待验证。

原文摘要 · Abstract (English)

Domain-camouflaged injection attacks embed malicious instructions in retrieved content using domain-appropriate vocabulary, evading standard detectors that rely on syntactic injection markers. When detection fails, practitioners need to know which defense architectures reduce attack success. We evaluate five prompting-based defenses (spotlighting, paraphrasing, prompt sandwiching, and two combinations) against domain-camouflaged injection across three model families (Claude Haiku, Llama 3.1 8B, Gemini 2.0 Flash) and three deployment domains (financial, legal, general) using 3,510 trials. Paraphrasing retrieved content before agent processing is the most consistently effective defense in this benchmark, reducing camouflage attack success rate by 55-84\% depending on model, and achieves lower attack success rates than our Llama Guard 4 configuration on every model tested. Defense effectiveness is strongly model-dependent: spotlighting halves attack success on Claude Haiku but provides no benefit on Llama 3.1 8B. Financial domain deployments face the highest residual risk at 26-33\% baseline attack success rate, with no prompting-based defense fully eliminating the threat on weaker models. These results provide the first systematic evaluation of prompting-based defenses specifically against camouflage-class injection attacks and establish benchmark-based recommendations for practitioners. All tasks use synthetically constructed professional documents; whether these benchmark rankings generalize to real enterprise documents remains an open question.

提示防御隐蔽攻击模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。