arXiv:2508.16484cs.CL2025-08被引 2

自动化生成隐蔽的越狱提示,突破紧凑大模型安全防护。

HAMSA: Hijacking Aligned Compact Models via Stealthy Automation

  • 多阶段进化搜索,用温度调控保持语言流畅性
  • 在英阿双语数据集上成功绕过对齐机制
  • 适合安全研究者与模型防御团队参考

大型语言模型,尤其是高效紧凑型变体,仍易受越狱攻击影响,即使经过大量对齐训练。现有对抗性提示生成方法多依赖人工设计或简单混淆,生成文本质量低且易被困惑度过滤器识别。本文提出一种自动化红队框架,可为对齐的紧凑大模型生成语义清晰、隐蔽性强的越狱提示。该方法采用多阶段进化搜索,通过种群策略结合温度控制的变异机制,在探索与语言连贯性之间取得平衡,系统发现能绕过对齐防护且保持自然语言流畅性的提示。我们在英文基准(In-The-Wild Jailbreak Prompts on LLMs)和新构建的阿拉伯语数据集(基于原英文数据并由母语阿拉伯语语言学家标注)上进行评估,实现多语言安全性测试。

原文摘要 · Abstract (English)

Large Language Models (LLMs), especially their compact efficiency-oriented variants, remain susceptible to jailbreak attacks that can elicit harmful outputs despite extensive alignment efforts. Existing adversarial prompt generation techniques often rely on manual engineering or rudimentary obfuscation, producing low-quality or incoherent text that is easily flagged by perplexity-based filters. We present an automated red-teaming framework that evolves semantically meaningful and stealthy jailbreak prompts for aligned compact LLMs. The approach employs a multi-stage evolutionary search, where candidate prompts are iteratively refined using a population-based strategy augmented with temperature-controlled variability to balance exploration and coherence preservation. This enables the systematic discovery of prompts capable of bypassing alignment safeguards while maintaining natural language fluency. We evaluate our method on benchmarks in English (In-The-Wild Jailbreak Prompts on LLMs), and a newly curated Arabic one derived from In-The-Wild Jailbreak Prompts on LLMs and annotated by native Arabic linguists, enabling multilingual assessment.

越狱攻击自动化红队多语言安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。