arXiv:2602.22242cs.CRcs.AI2026-02被引 2

测试多个开源大模型在提示注入攻击下的安全表现

Analysis of LLMs Against Prompt Injection and Jailbreak Attacks

  • 构建人工标注数据集评估提示注入漏洞
  • 不同模型响应差异大,部分完全无响应
  • 轻量级防御可被复杂推理提示绕过

大型语言模型(LLMs)广泛部署于实际系统中。随着其应用范围扩大,提示工程已成为资源有限组织使用 LLM 的有效手段。然而,LLMs 易受基于提示的攻击影响,因此分析此类风险已成为关键安全需求。本文利用大规模人工标注数据集,在多个开源 LLM 上评估了提示注入与越狱攻击的脆弱性,包括 Phi、Mistral、DeepSeek-R1、Llama 3.2、Qwen 及 Gemma 变体。我们观察到模型间行为差异显著,包括由内部安全机制触发的拒绝响应与完全沉默无响应。此外,我们评估了几种轻量级、无需重训练或高算力微调的推理时防御机制。尽管这些防御可缓解简单攻击,但均被长篇、高推理密度的提示持续绕过。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are widely deployed in real-world systems. Given their broader applicability, prompt engineering has become an efficient tool for resource-scarce organizations to adopt LLMs for their own purposes. At the same time, LLMs are vulnerable to prompt-based attacks. Thus, analyzing this risk has become a critical security requirement. This work evaluates prompt-injection and jailbreak vulnerability using a large, manually curated dataset across multiple open-source LLMs, including Phi, Mistral, DeepSeek-R1, Llama 3.2, Qwen, and Gemma variants. We observe significant behavioural variation across models, including refusal responses and complete silent non-responsiveness triggered by internal safety mechanisms. Furthermore, we evaluated several lightweight, inference-time defence mechanisms that operate as filters without any retraining or GPU-intensive fine-tuning. Although these defences mitigate straightforward attacks, they are consistently bypassed by long, reasoning-heavy prompts.

大模型安全提示攻击防御机制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。