发现大模型在不同语言和时间表述下安全表现剧烈波动,暴露全球南方用户风险。
Safe in the Future, Dangerous in the Past: Dissecting Temporal and Linguistic Vulnerabilities in LLMs
- 构建西非威胁场景的对抗数据集,测试多语言与时间框架下的安全表现
- 未来时态触发过度拒绝(57.2%安全),过去时态则防线崩溃(15.6%安全)
- 模型安全依赖表面启发式,适合关注AI公平性与系统鲁棒性的研究者
随着大语言模型融入全球关键基础设施,人们普遍认为英语的安全对齐可零样本迁移至其他语言,这一假设实为危险盲区。本研究对三款顶尖模型(GPT-5.1、Gemini 3 Pro、Claude 4.5 Opus)进行系统审计,采用基于西非威胁场景(如Yahoo-Yahoo诈骗、Dane枪支制造)的HausaSafety对抗数据集,通过1,440次评估的2×4因子设计,检验语言(英语 vs. Hausa)与时间表述的非线性交互。结果挑战了多语言安全差距的简单叙事:并非低资源语言普遍脆弱,而是存在复杂干扰机制。虽Claude 4.5 Opus在豪萨语中更安全(45.0%),高于英语(36.7%),源于不确定性驱动拒绝,但在时间推理上遭遇灾难性失败。报告显著的时间不对称性:过去时表述使防御失效(仅15.6%安全),而未来时触发超保守拒绝(57.2%安全)。最极端配置间安全率相差9.2倍,证明安全并非固有属性,而是情境依赖状态。结论指出当前模型依赖表面启发式,缺乏深层语义理解,形成‘安全孤岛’,使全球南方用户暴露于局部危害。提出不变对齐(Invariant Alignment)作为必要范式转变,以保障跨语言与时间变化中的安全稳定性。
原文摘要 · Abstract (English)
As Large Language Models (LLMs) integrate into critical global infrastructure, the assumption that safety alignment transfers zero-shot from English to other languages remains a dangerous blind spot. This study presents a systematic audit of three state of the art models (GPT-5.1, Gemini 3 Pro, and Claude 4.5 Opus) using HausaSafety, a novel adversarial dataset grounded in West African threat scenarios (e.g., Yahoo-Yahoo fraud, Dane gun manufacturing). Employing a 2 x 4 factorial design across 1,440 evaluations, we tested the non-linear interaction between language (English vs. Hausa) and temporal framing. Our results challenge the narrative of the multilingual safety gap. Instead of a simple degradation in low-resource settings, we identified a complex interference mechanism in which safety is determined by the intersection of variables. Although the models exhibited a reverse linguistic vulnerability with Claude 4.5 Opus proving significantly safer in Hausa (45.0%) than in English (36.7%) due to uncertainty-driven refusal, they suffered catastrophic failures in temporal reasoning. We report a profound Temporal Asymmetry, where past-tense framing bypassed defenses (15.6% safe) while future-tense scenarios triggered hyper-conservative refusals (57.2% safe). The magnitude of this volatility is illustrated by a 9.2x disparity between the safest and most vulnerable configurations, proving that safety is not a fixed property but a context-dependent state. We conclude that current models rely on superficial heuristics rather than robust semantic understanding, creating Safety Pockets that leave Global South users exposed to localized harms. We propose Invariant Alignment as a necessary paradigm shift to ensure safety stability across linguistic and temporal shifts.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。