arXiv:2507.22133cs.CRcs.CL2025-07被引 2

优化提示词提升大模型攻击生成效率,发现隐藏漏洞模式。

Prompt Optimization and Evaluation for LLM Automated Red Teaming

  • 用成功率评估单次攻击的可发现性,动态优化提示词。
  • 重复攻击测试揭示高成功率攻击模式,显著提升漏洞探测能力。
  • 适合安全研究者与大模型防御系统开发者参考。

大型语言模型(LLM)应用日益广泛,其系统漏洞识别愈发重要。自动化红队通过使用 LLM 生成并执行攻击来加速漏洞探测。攻击生成器的性能通常以攻击成功率(ASR)衡量,即对每轮攻击成功判断的样本均值。本文提出一种基于单个攻击成功率的新方法,通过在随机种子的目标上重复执行攻击,测量攻击的可发现性——即单次攻击的成功期望值。该方法揭示了可被利用的攻击模式,指导提示词优化,从而实现更稳健的生成器评估与迭代改进。

原文摘要 · Abstract (English)

Applications that use Large Language Models (LLMs) are becoming widespread, making the identification of system vulnerabilities increasingly important. Automated Red Teaming accelerates this effort by using an LLM to generate and execute attacks against target systems. Attack generators are evaluated using the Attack Success Rate (ASR) the sample mean calculated over the judgment of success for each attack. In this paper, we introduce a method for optimizing attack generator prompts that applies ASR to individual attacks. By repeating each attack multiple times against a randomly seeded target, we measure an attack's discoverability the expectation of the individual attack success. This approach reveals exploitable patterns that inform prompt optimization, ultimately enabling more robust evaluation and refinement of generators.

大模型安全红队攻击提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。