arXiv:2505.12287cs.CLcs.AI2025-05被引 2

首次系统评估闭源大模型多语言越狱攻击,发现中文攻击更有效。

The Tower of Babel Revisited: Multilingual Jailbreak Prompts on Closed-Source Large Language Models

  • 构建多语言越狱攻击框架,覆盖32种攻击类型
  • 测试4款闭源模型,中文攻击成功率普遍高于英文
  • 提出双面攻击法,对所有模型均最有效

大型语言模型(LLMs)广泛应用,但易受对抗性提示注入攻击。现有研究多聚焦开源模型,本文首次系统评估闭源模型在多语言场景下的安全弱点。构建集成攻击框架,测试GPT-4o、DeepSeek-R1、Gemini-1.5-Pro和Qwen-Max四款前沿模型,覆盖中英文6类安全内容,生成38,400条响应。采用攻击成功率(ASR)从提示设计、模型架构、语言环境三维度评估。结果表明:Qwen-Max最易受攻击,GPT-4o防御最强;中文提示的平均攻击成功率显著高于英文;新提出的两面攻击法在所有模型上表现最优。研究揭示了语言感知对齐与跨语言防御的紧迫需求,呼吁推动更鲁棒、包容的AI系统发展。

原文摘要 · Abstract (English)

Large language models (LLMs) have seen widespread applications across various domains, yet remain vulnerable to adversarial prompt injections. While most existing research on jailbreak attacks and hallucination phenomena has focused primarily on open-source models, we investigate the frontier of closed-source LLMs under multilingual attack scenarios. We present a first-of-its-kind integrated adversarial framework that leverages diverse attack techniques to systematically evaluate frontier proprietary solutions, including GPT-4o, DeepSeek-R1, Gemini-1.5-Pro, and Qwen-Max. Our evaluation spans six categories of security contents in both English and Chinese, generating 38,400 responses across 32 types of jailbreak attacks. Attack success rate (ASR) is utilized as the quantitative metric to assess performance from three dimensions: prompt design, model architecture, and language environment. Our findings suggest that Qwen-Max is the most vulnerable, while GPT-4o shows the strongest defense. Notably, prompts in Chinese consistently yield higher ASRs than their English counterparts, and our novel Two-Sides attack technique proves to be the most effective across all models. This work highlights a dire need for language-aware alignment and robust cross-lingual defenses in LLMs, and we hope it will inspire researchers, developers, and policymakers toward more robust and inclusive AI systems.

越狱攻击多语言闭源模型安全评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。