arXiv:2511.00689cs.CL2025-11被引 5

首次系统评估多语言下大模型越狱攻击与防御的泛化能力

Do Methods to Jailbreak and Defend LLMs Generalize Across Languages?

  • 在10种语言上测试六种大模型的越狱攻击与防御效果
  • 高资源语言对常规查询更安全,但对对抗性提示更易受攻击
  • 简单防御策略效果依赖语言和模型,需语言敏感的安全评估

大型语言模型(LLMs)在训练后经过安全对齐,但近期研究表明其安全机制可被越狱攻击绕过。尽管已有多种越狱方法和防御措施,其跨语言泛化性仍缺乏研究。本文首次在十种语言(涵盖高、中、低资源语言)上,基于HarmBench和AdvBench基准,系统评估六种大模型在两类越狱攻击下的表现:基于逻辑表达和对抗性提示的攻击。结果表明,攻击成功率与防御鲁棒性均随语言变化;高资源语言在标准查询下更安全,但在对抗性提示下更脆弱。简单的防御策略虽有效,但具有语言和模型依赖性。研究呼吁构建语言感知且具备跨语言覆盖能力的大模型安全评估基准。

原文摘要 · Abstract (English)

Large language models (LLMs) undergo safety alignment after training and tuning, yet recent work shows that safety can be bypassed through jailbreak attacks. While many jailbreaks and defenses exist, their cross-lingual generalization remains underexplored. This paper presents the first systematic multilingual evaluation of jailbreaks and defenses across ten languages -- spanning high-, medium-, and low-resource languages -- using six LLMs on HarmBench and AdvBench. We assess two jailbreak types: logical-expression-based and adversarial-prompt-based. For both types, attack success and defense robustness vary across languages: high-resource languages are safer under standard queries but more vulnerable to adversarial ones. Simple defenses can be effective, but are language- and model-dependent. These findings call for language-aware and cross-lingual safety benchmarks for LLMs.

大模型安全越狱攻击多语言评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。