arXiv:2506.22557cs.CRcs.LG2025-06AAAI被引 2

用强化学习打造低成本通用攻击框架,10次查询即突破大模型安全防线

MetaCipher: A Time-Persistent and Universal Multi-Agent Framework for Cipher-Based Jailbreak Attacks for LLMs

  • 构建多智能体系统,通过强化学习实现跨模型通用攻击
  • 仅需10次查询即在最新恶意提示基准上达顶尖成功率
  • 适合研究对抗攻击与模型安全的学者,警示高风险内容

随着大语言模型能力提升,其面临日益复杂的越狱攻击威胁。尽管开发者投入大量资源进行对齐微调和安全防护,研究者仍不断提出新攻击方法,推动攻防技术持续迭代。然而,两大挑战限制了越狱研究的效率与影响:顶级模型查询成本高昂,且有效攻击策略因频繁安全更新而寿命短。为此,我们提出MetaCipher,一种低代价、可泛化的多智能体越狱框架,能适应具有不同安全机制的多个大模型。该框架基于强化学习,模块化设计且具备可扩展性,支持未来策略演进。实验表明,仅需10次查询,即可在近期恶意提示基准上达到最先进的攻击成功率,显著优于现有方法。我们在多种目标模型与基准上开展大规模实证评估,验证了其鲁棒性与适应性。警告:本文包含可能具冒犯性或有害的模型输出,仅用于展示越狱有效性。

原文摘要 · Abstract (English)

As large language models (LLMs) grow more capable, they face growing vulnerability to sophisticated jailbreak attacks. While developers invest heavily in alignment finetuning and safety guardrails, researchers continue publishing novel attacks, driving progress through adversarial iteration. This dynamic mirrors a strategic game of continual evolution. However, two major challenges hinder jailbreak development: the high cost of querying top-tier LLMs and the short lifespan of effective attacks due to frequent safety updates. These factors limit cost-efficiency and practical impact of research in jailbreak attacks. To address this, we propose MetaCipher, a low-cost, multi-agent jailbreak framework that generalizes across LLMs with varying safety measures. Using reinforcement learning, MetaCipher is modular and adaptive, supporting extensibility to future strategies. Within as few as 10 queries, MetaCipher achieves state-of-the-art attack success rates on recent malicious prompt benchmarks, outperforming prior jailbreak methods. We conduct a large-scale empirical evaluation across diverse victim models and benchmarks, demonstrating its robustness and adaptability. Warning: This paper contains model outputs that may be offensive or harmful, shown solely to demonstrate jailbreak efficacy.

越狱攻击强化学习多智能体大模型安全

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。