用模块化组合攻击提升大模型越狱成功率,效率更高且更易扩展。
MAD-MAX: Modular And Diverse Malicious Attack MiXtures for Automated LLM Red Teaming
- 将攻击策略分组后智能组合,生成多样化新攻击
- 在GPT-4o和Gemini-Pro上实现97%越狱成功率,仅需10.9次查询
- 可自动融合新攻击方法,适合持续安全测试场景
随着大语言模型应用日益广泛,其易受越狱攻击导致有害输出的问题构成重大安全风险。新越狱策略不断涌现,模型亦频繁通过微调更新,因此持续进行漏洞检测至关重要。现有红队测试方法在成本效率、攻击成功率、多样性或可扩展性方面存在不足。本文提出模块化且多样化的恶意攻击混合方法(MAD-MAX),实现自动化大模型红队测试。MAD-MAX通过自动将攻击策略分配至相关攻击簇,根据恶意目标选择最相关的簇,并从中组合策略以生成高成功率、多样化的新型攻击。该方法在每次红队迭代中融合有前景的攻击以提升性能,并引入相似性过滤机制剔除重复攻击,提高成本效率。MAD-MAX设计为可轻松集成新发现的攻击策略,在攻击成功率(ASR)与所需查询次数上显著优于主流方法Tree of Attacks with Pruning(TAP)。在基准测试中,MAD-MAX对GPT-4o和Gemini-Pro实现97%的恶意目标越狱成功率,远超TAP的66%;且平均仅需10.9次查询,低于TAP的23.3次。警告:本文包含具有冒犯性的内容。
原文摘要 · Abstract (English)
With LLM usage rapidly increasing, their vulnerability to jailbreaks that create harmful outputs are a major security risk. As new jailbreaking strategies emerge and models are changed by fine-tuning, continuous testing for security vulnerabilities is necessary. Existing Red Teaming methods fall short in cost efficiency, attack success rate, attack diversity, or extensibility as new attack types emerge. We address these challenges with Modular And Diverse Malicious Attack MiXtures (MAD-MAX) for Automated LLM Red Teaming. MAD-MAX uses automatic assignment of attack strategies into relevant attack clusters, chooses the most relevant clusters for a malicious goal, and then combines strategies from the selected clusters to achieve diverse novel attacks with high attack success rates. MAD-MAX further merges promising attacks together at each iteration of Red Teaming to boost performance and introduces a similarity filter to prune out similar attacks for increased cost efficiency. The MAD-MAX approach is designed to be easily extensible with newly discovered attack strategies and outperforms the prominent Red Teaming method Tree of Attacks with Pruning (TAP) significantly in terms of Attack Success Rate (ASR) and queries needed to achieve jailbreaks. MAD-MAX jailbreaks 97% of malicious goals in our benchmarks on GPT-4o and Gemini-Pro compared to TAP with 66%. MAD-MAX does so with only 10.9 average queries to the target LLM compared to TAP with 23.3. WARNING: This paper contains contents which are offensive in nature.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。