arXiv:2605.14152cs.CLcs.AI2026-05

通过跨文化对抗测试,揭示语言与地缘政治如何共同影响大模型安全行为。

ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety

论文配图:ROK-FORTRESS: Measuring the Effect of Geopolitical Transcreation for National Security and Public Safety
图 1 · 摘自论文原文
  • 构建英文-韩文对抗性评测矩阵,分离语言与地缘背景影响
  • 韩语版本模型普遍存在安全响应抑制,且受地缘背景调节
  • 适合关注多语言安全评估与地缘敏感模型的开发者

大型语言模型(LLM)在国家安全与公共安全(NSPS)风险评估中日益重要,但多语言安全评估仍主要依赖翻译基准,未充分考察语言与地缘政治背景的交互作用。本文提出ROK-FORTRESS,一个以美韩关系为案例的双语、文化对抗性NSPS评测基准,采用英-韩语言对和美-韩地缘轴心,通过转创矩阵(transcreation matrix)控制变量:分别测试(i)英语与韩语语言形式,以及(ii)美国与韩国实体、机构和操作细节的组合。每条恶意提示均配有一条双重用途良性对照提示,用于量化过度拒绝现象;响应由经校准的LLM评委组根据专家设计的、提示特定的二元评分标准打分。在前沿模型与韩语优化模型组成的双轨测试中,发现韩语变体存在持续的抑制效应,且模型间地缘背景与语言交互差异显著;部分模型中,韩语地缘背景进一步削弱了语言驱动的抑制。这表明,在英-韩案例中,安全行为受语言作为风险信号及上下文交互的影响,而翻译仅评估无法捕捉此类机制。直接请求消融实验显示,闭源模型有轻微但持续的降低,而开源模型则出现更大且依赖提示封装的反转效应,暗示部分韩语抑制源于提示特化而非内在语言安全对齐。该转创矩阵方法可推广至其他语言-文化对。

原文摘要 · Abstract (English)

Safety evaluations for large language models (LLMs) increasingly target high-stakes National Security and Public Safety (NSPS) risks, yet multilingual safety is mostly assessed through translation-only benchmarks that preserve the underlying scenario, leaving how language and geopolitical context interact largely unexamined beyond a few language pairs. We introduce ROK-FORTRESS, a bilingual, culturally adversarial NSPS benchmark that uses the English-Korean language pair and U.S.-ROK geopolitical axis as a case study, separating the effects of language and geopolitical grounding via a transcreation matrix: adversarial intents are evaluated under controlled combinations of (i) English versus Korean language and (ii) U.S. versus Korean entities, institutions, and operational details. Each adversarial prompt is paired with a dual-use benign counterpart to quantify over-refusal, and responses are scored by calibrated LLM-as-a-judge panels using expert-crafted, prompt-specific binary rubrics. Across a dual-track set of frontier and Korean-optimized models, we find a consistent suppression effect in Korean variants and substantial model-to-model variation in how geopolitical grounding interacts with language; in a subset of models, Korean grounding further mitigates the language-driven suppression. This indicates that, at least in the English-Korean case, safety behavior is shaped by language-as-risk signals and context interactions that translation-only evaluations miss. A direct-request ablation that strips jailbreak wrappers separates a small but persistent reduction for closed-source models from a larger, wrapper-dependent effect that reverses for open-source models, suggesting part of the Korean suppression reflects prompt specialization rather than intrinsic language-based safety alignment. The transcreation matrix methodology is designed to generalize to other language-culture pairs.

多语言安全地缘政治模型评测

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。