多策略协同攻击,突破大模型安全防线。
Transferable & Stealthy Ensemble Attacks: A Black-Box Jailbreaking Framework for Large Language Models
- 融合多种攻击策略形成集成攻击
- 在2024年安全竞赛中获攻击赛道第一
- 适合研究模型漏洞与对抗攻击者
我们提出一种新型黑盒越狱框架,整合多种大模型作为攻击者的策略,实现高度可迁移且高效的攻击。该框架基于三个关键洞察:集成方法比单一方法更能暴露对齐大模型的漏洞;恶意指令的越狱难度各异,需针对性优化;破坏恶意提示的语义连贯性可操纵其嵌入表示,提升攻击成功率。在2024年大模型与智能体安全竞赛中,该方案在越狱攻击赛道取得第一名。
原文摘要 · Abstract (English)
We present a novel black-box jailbreaking framework that integrates multiple LLM-as-Attacker strategies to deliver highly transferable and effective attacks. The framework is grounded in three key insights from prior jailbreaking research and practice: ensemble approaches outperform single methods in exposing aligned LLM vulnerabilities, malicious instructions vary in jailbreaking difficulty requiring tailored optimization, and disrupting semantic coherence of malicious prompts can manipulate their embeddings to boost success rates. Validated in the Competition for LLM and Agent Safety 2024, our solution achieved top rankings in the Jailbreaking Attack Track.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。