arXiv:2601.04666cs.AIcs.CR2026-01ACL被引 1

通过合成多样数据与指令级思维链训练,提升大模型对提示注入攻击的防御能力

Know Thy Enemy: Securing LLMs Against Prompt Injection via Diverse Data Synthesis and Instruction-Level Chain-of-Thought Learning

  • 合成多样化恶意指令数据,增强模型对复杂攻击的识别能力
  • 在四种大模型上验证,对行为偏移、隐私泄露和有害输出防护效果显著提升
  • 适合安全敏感场景下的大模型应用,如金融、医疗对话系统

集成大语言模型(LLM)的应用日益广泛,但面临提示注入(PI)攻击的严重安全威胁。现有防御方法存在两大难题:恶意指令可通过多种途径注入,且注入内容常与上下文语义边界模糊,难以识别。为此,我们提出InstruCoT,一种面向PI防御的模型增强方法,通过合成多样化训练数据并采用指令级思维链微调,使LLM能有效识别并拒绝任意来源或位置的恶意指令。我们在行为偏差、隐私泄露和有害输出三个关键维度对InstruCoT进行了评估。实验结果表明,在四种主流大模型上,InstruCoT在所有维度均显著优于基线方法,同时保持原有任务性能未下降。

原文摘要 · Abstract (English)

Large language model (LLM)-integrated applications have become increasingly prevalent, yet face critical security vulnerabilities from prompt injection (PI) attacks. Defending against PI attacks faces two major issues: malicious instructions can be injected through diverse vectors, and injected instructions often lack clear semantic boundaries from the surrounding context, making them difficult to identify. To address these issues, we propose InstruCoT, a model enhancement method for PI defense that synthesizes diverse training data and employs instruction-level chain-of-thought fine-tuning, enabling LLMs to effectively identify and reject malicious instructions regardless of their source or position in the context. We evaluate InstruCoT across three critical dimensions: Behavior Deviation, Privacy Leakage, and Harmful Output. Experimental results across four LLMs demonstrate that InstruCoT significantly outperforms baselines in all dimensions while maintaining utility performance without degradation

提示注入大模型安全防御机制思维链

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。