自动化发现长尾分布下的大模型越狱攻击,提升安全评估效率
Evolving Jailbreaks: Automated Multi-Objective Long-Tail Attacks on Large Language Models
- 用多目标进化算法自动生成越狱提示,兼顾攻击效果与输出自然度
- 在多个长尾场景中成功触发越狱,性能媲美现有方法
- 适合关注大模型安全、对抗攻击研究者参考
大型语言模型(LLMs)广泛部署于开放的网络应用中,面临来自长尾分布(如低资源语言、加密私有数据)用户输入的安全风险。现有越狱攻击多依赖人工规则,难以系统评估漏洞。本文提出EvoJail框架,通过多目标进化搜索自动发现长尾攻击。将攻击提示生成建模为最大化攻击效果、最小化输出困惑度的优化问题,采用语义-算法混合表示捕捉意图与加密逻辑变换。结合LLM辅助的变异与交叉算子,在高度结构化的开放空间中高效探索。大量实验表明,EvoJail持续发现多样且有效的长尾越狱策略,在个体与集成层面表现均优于或相当现有方法。
原文摘要 · Abstract (English)
Large Language Models (LLMs) have been widely deployed, especially through free Web-based applications that expose them to diverse user-generated inputs, including those from long-tail distributions such as low-resource languages and encrypted private data. This open-ended exposure increases the risk of jailbreak attacks that undermine model safety alignment. While recent studies have shown that leveraging long-tail distributions can facilitate such jailbreaks, existing approaches largely rely on handcrafted rules, limiting the systematic evaluation of these security and privacy vulnerabilities. In this work, we present EvoJail, an automated framework for discovering long-tail distribution attacks via multi-objective evolutionary search. EvoJail formulates long-tail attack prompt generation as a multi-objective optimization problem that jointly maximizes attack effectiveness and minimizes output perplexity, and introduces a semantic-algorithmic solution representation to capture both high-level semantic intent and low-level structural transformations of encryption-decryption logic. Building upon this representation, EvoJail integrates LLM-assisted operators into a multi-objective evolutionary framework, enabling adaptive and semantically informed mutation and crossover for efficiently exploring a highly structured and open-ended search space. Extensive experiments demonstrate that EvoJail consistently discovers diverse and effective long-tail jailbreak strategies, achieving competitive performance with existing methods in both individual and ensemble level.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。