arXiv:2605.17173cs.CLcs.AI2026-05

揭示多语言安全防护失效的深层原因,提出可分解的评估框架。

Why Do Safety Guardrails Degrade Across Languages?

论文配图:Why Do Safety Guardrails Degrade Across Languages?
图 1 · 摘自论文原文
  • 用多组项目反应理论拆解安全性的语言无关鲁棒性、提示难度等四类因素。
  • 发现22个模型在英语中比低资源语言更易被越狱,且低资源语言响应熵更高。
  • 框架能预测跨语言安全表现,适合用于公平评估和改进多语言安全数据集。

大型语言模型在非英语语言中表现出安全性能下降。标准评估依赖越狱成功率(JSR),将多种安全驱动因素混为一谈,掩盖了真实成因。本文引入一种潜在变量模型——多组项目反应理论(Multi-Group IRT)框架,分离出语言无关的安全鲁棒性(θ)、提示固有难度(β)、全球语言处理难度(γ)以及提示特定的跨语言安全差距(τ)。基于MultiJail数据集,评估了5个闭源模型家族共61种配置在10种不同资源语言中的表现,累计生成190万条响应。探索性因子分析显示安全行为具有高度单维性:模型拒绝各类危害主要依赖同一机制。与预期相反,22种模型配置在英语中比在低资源语言中更脆弱;低资源语言产生更高熵的不确定回应。高τ提示集中于盗窃、武器等物理危害类别,并与低资源语言相关,该趋势经跨数据集验证。尽管整体翻译质量与τ相关性低,但严重误译会引发高偏倚异常值,经母语者验证确认。文化与概念基础不匹配也可能贡献τ。预测验证中,IRT框架达到AUC=0.940,且在整门语言被排除时仍保持预测力(0.875),优于率基线。该框架揭示了聚合指标所掩盖的概念-语言漏洞,有助于实现更公平的跨语言安全评估与数据集优化。

原文摘要 · Abstract (English)

Large language models exhibit safety degradation in non-English languages. Standard evaluation relies on Jailbreak Success Rate (JSR), which confounds several safety-driving factors into one, obscuring the specific cause(s) of safety failure. We introduce a latent variable model, a Multi-Group Item Response Theory (IRT) framework, that decouples language-agnostic safety robustness ($θ$), intrinsic prompt hardness ($β$), global language processing difficulty ($γ$), and a prompt-specific cross-lingual safety gap ($τ$). Using the MultiJail dataset, we evaluate the safety robustness of 61 model configurations across 5 closed-model families and 10 languages of varying resource, aggregating a dataset of 1.9 million responses. Exploratory Factor Analysis shows safety is primarily unidimensional: models refuse different harm types mainly through a shared mechanism. Contrary to the expected trend that safety degrades largely in low-resource languages, 22 model configurations are more vulnerable in English than in low-resource languages. Low-resource languages produce more uncertain responses (high entropy) than high-resource languages. Also, high-$τ$ prompts cluster in physical harm categories like Theft and Weapons and lower-resource languages, trends validated through cross-dataset generalization. While global translation quality shows low correlation with $τ$, severe mistranslations drive high-bias outliers, as validated by native speakers. Cultural and conceptual grounding mismatches may also contribute to $τ$. In predictive validation, the IRT framework achieves $\mathrm{AUC} = 0.940$, and unlike rate baselines stays predictive when a whole language is held out ($0.875$). Our framework reveals concept-language vulnerabilities that aggregate metrics obscure, enabling fairer cross-lingual safety evaluation and targeted improvements in dataset construction.

多语言安全模型评估项目反应理论越狱防御

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。