arXiv:2605.05662cs.CLcs.AI2026-05被引 2

构建跨文化安全评测基准,精准评估大模型在不同国家的敏感性识别能力。

XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity

论文配图:XL-SafetyBench: A Country-Grounded Cross-Cultural Benchmark for LLM Safety and Cultural Sensitivity
图 1 · 摘自论文原文
  • 设计多阶段流程生成5500个跨国家语言测试用例,包含对抗性攻击与文化敏感请求。
  • 发现顶级模型在越狱防御与文化敏感度上无关联,本地模型存在安全假象。
  • 引入新指标区分拒绝行为与理解失败,适合多语言安全评估研究者使用。

当前的大模型安全评测主要以英语为中心,依赖翻译,难以捕捉国家特有的危害。且很少评估模型对文化嵌入敏感性的识别能力。我们提出XL-SafetyBench,涵盖10个语种-国家组合的5500个测试用例,包括基于国家背景的越狱攻击基准和嵌入本土敏感性的文化敏感性基准。每项通过多阶段流程生成:大模型辅助发现、自动化验证及双独立母语标注员。为区分合理拒绝与理解失败,引入攻击成功率(ASR)、中性安全率(NSR)与文化敏感率(CSR)三项指标。评估10个前沿模型与27个本地模型发现:第一,顶级模型在越狱鲁棒性与文化意识间无耦合关系,综合安全评分掩盖了维度差异;第二,本地模型呈现近似线性的ASR-NSR权衡(r = -0.81),表明其表面安全性实为生成失败而非真实对齐。该基准支持多语言时代更精细的跨文化安全评估。

原文摘要 · Abstract (English)

Current LLM safety benchmarks are predominantly English-centric and often rely on translation, failing to capture country-specific harms. Moreover, they rarely evaluate a model's ability to detect culturally embedded sensitivities as distinct from universal harms. We introduce XL-SafetyBench. a suite of 5,500 test cases across 10 country-language pairs, comprising a Jailbreak Benchmark of country-grounded adversarial prompts and a Cultural Benchmark where local sensitivities are embedded within innocuous requests. Each item is constructed via a multi-stage pipeline that combines LLM-assisted discovery, automated validation gates, and dual independent native-speaker annotators per country. To distinguish principled refusal from comprehension failure, we evaluate Attack Success Rate (ASR) alongside two complementary metrics we introduce: Neutral-Safe Rate (NSR) and Cultural Sensitivity Rate (CSR). Evaluating 10 frontier and 27 local LLMs reveals two key findings. First, jailbreak robustness and cultural awareness do not show a coupled relationship among frontier models, so a composite safety score obscures per-axis variation. Second, local models exhibit a near-linear ASR-NSR trade-off (r = -0.81), indicating that their apparent safety reflects generation failure rather than genuine alignment. XL-SafetyBench enables more nuanced, cross-cultural safety evaluation in the multilingual era.

大模型安全跨文化评估评测基准

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。