自动演化提示,高效挖掘大模型安全漏洞。
Be a Multitude to Itself: A Prompt Evolution Framework for Red Teaming
- 通过广度与深度演化生成多样化攻击提示
- 在8个敏感话题上对8个大模型测试,成功率更高
- 适合安全研究者、模型开发者用于风险评估
大型语言模型(LLMs)虽能力强大,但存在生成有害内容的安全隐患。红队测试旨在发现能诱发有害响应的提示,是提前发现并缓解安全风险的关键。然而,人工红队测试耗时费力,难以扩展。本文提出RTPE框架,实现提示在广度和深度上的可扩展演化,自动生成大量高质量、多样化的红队提示。广度演化采用新型上下文学习方法生成多类优质提示;深度演化则通过定制化变换操作,增强提示的内容与形式多样性。大量实验表明,RTPE在攻击成功率和提示多样性上均优于现有主流自动红队方法。此外,基于RTPE生成的4,800个提示,我们对8个代表性大模型在8个敏感话题上进行了系统分析。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have gained increasing attention for their remarkable capacity, alongside concerns about safety arising from their potential to produce harmful content. Red teaming aims to find prompts that could elicit harmful responses from LLMs, and is essential to discover and mitigate safety risks before real-world deployment. However, manual red teaming is both time-consuming and expensive, rendering it unscalable. In this paper, we propose RTPE, a scalable evolution framework to evolve red teaming prompts across both breadth and depth dimensions, facilitating the automatic generation of numerous high-quality and diverse red teaming prompts. Specifically, in-breadth evolving employs a novel enhanced in-context learning method to create a multitude of quality prompts, whereas in-depth evolving applies customized transformation operations to enhance both content and form of prompts, thereby increasing diversity. Extensive experiments demonstrate that RTPE surpasses existing representative automatic red teaming methods on both attack success rate and diversity. In addition, based on 4,800 red teaming prompts created by RTPE, we further provide a systematic analysis of 8 representative LLMs across 8 sensitive topics.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。