arXiv:2505.00843cs.CRcs.AI2025-05被引 3

OET工具箱可自动生成对抗性提示,评估大模型防御能力。

OET: Optimization-based prompt injection Evaluation Toolkit

  • 基于优化方法生成自适应对抗提示,模拟真实攻击场景。
  • 测试显示部分强化模型仍易受攻击,暴露现有防御短板。
  • 适合安全研究人员和模型开发者用于严格红队测试。

大型语言模型在自然语言理解与生成方面表现出色,但其易受提示注入攻击影响,可能导致行为被操纵或指令被覆盖。尽管已有多种防御策略,却缺乏标准框架来系统评估其在自适应攻击下的有效性。为此,我们提出OET——一种基于优化的提示注入评估工具包,通过自适应测试框架,在多种数据集上系统化地评测攻击与防御效果。该工具包具备模块化工作流,支持对抗性字符串生成、动态攻击执行与全面结果分析,提供统一平台以评估模型的对抗鲁棒性。关键在于,自适应测试框架结合白盒与黑盒访问,利用优化方法生成最坏情况的对抗样本,实现严格的红队测试。大量实验表明,当前防御机制存在明显局限,部分模型即使经过安全增强依然脆弱。

原文摘要 · Abstract (English)

Large Language Models (LLMs) have demonstrated remarkable capabilities in natural language understanding and generation, enabling their widespread adoption across various domains. However, their susceptibility to prompt injection attacks poses significant security risks, as adversarial inputs can manipulate model behavior and override intended instructions. Despite numerous defense strategies, a standardized framework to rigorously evaluate their effectiveness, especially under adaptive adversarial scenarios, is lacking. To address this gap, we introduce OET, an optimization-based evaluation toolkit that systematically benchmarks prompt injection attacks and defenses across diverse datasets using an adaptive testing framework. Our toolkit features a modular workflow that facilitates adversarial string generation, dynamic attack execution, and comprehensive result analysis, offering a unified platform for assessing adversarial robustness. Crucially, the adaptive testing framework leverages optimization methods with both white-box and black-box access to generate worst-case adversarial examples, thereby enabling strict red-teaming evaluations. Extensive experiments underscore the limitations of current defense mechanisms, with some models remaining susceptible even after implementing security enhancements.

提示攻击安全评估对抗样本LLM

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。