动态生成攻击提示,全面测试大模型安全防护能力
GuardVal: Dynamic Large Language Model Jailbreak Evaluation for Comprehensive Safety Testing
- 根据防御模型状态动态生成并优化攻击提示
- 在10个安全领域测试从Mistral-7b到GPT-4的模型表现
- 揭示模型行为差异,助力构建更安全的AI系统
越狱攻击暴露了大语言模型(LLMs)在生成有害或不道德内容方面的关键漏洞。由于模型持续演进且攻击手段日益复杂,评估这些威胁极具挑战性。现有基准和评估方法难以全面应对,导致对模型漏洞的评估存在缺口。本文回顾现有越狱评估实践,识别出有效评估协议的三个理想特性。为此,我们提出GuardVal,一种动态生成并基于防御模型状态优化越狱提示的新评估协议,可更准确评估防御模型在安全关键场景下的应对能力。此外,我们设计了一种新优化方法,防止提示优化过程停滞,确保生成越来越有效的攻击提示,从而暴露防御模型更深层的弱点。我们将该协议应用于涵盖从Mistral-7b到GPT-4的多种模型,在10个安全领域进行测试。结果揭示了各模型的显著行为差异,提供了其鲁棒性的全景视图。同时,评估过程深化了对模型行为的理解,为未来研究提供洞见,推动更安全模型的发展。
原文摘要 · Abstract (English)
Jailbreak attacks reveal critical vulnerabilities in Large Language Models (LLMs) by causing them to generate harmful or unethical content. Evaluating these threats is particularly challenging due to the evolving nature of LLMs and the sophistication required in effectively probing their vulnerabilities. Current benchmarks and evaluation methods struggle to fully address these challenges, leaving gaps in the assessment of LLM vulnerabilities. In this paper, we review existing jailbreak evaluation practices and identify three assumed desiderata for an effective jailbreak evaluation protocol. To address these challenges, we introduce GuardVal, a new evaluation protocol that dynamically generates and refines jailbreak prompts based on the defender LLM's state, providing a more accurate assessment of defender LLMs' capacity to handle safety-critical situations. Moreover, we propose a new optimization method that prevents stagnation during prompt refinement, ensuring the generation of increasingly effective jailbreak prompts that expose deeper weaknesses in the defender LLMs. We apply this protocol to a diverse set of models, from Mistral-7b to GPT-4, across 10 safety domains. Our findings highlight distinct behavioral patterns among the models, offering a comprehensive view of their robustness. Furthermore, our evaluation process deepens the understanding of LLM behavior, leading to insights that can inform future research and drive the development of more secure models.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。