arXiv:2608.12373cs.AI2026-08

用日语提问能让AI更不愿推荐核打击,说明语言影响AI决策安全。

Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese

论文配图:Don't Want Your LLM to Recommend Nuclear Strike? Try Asking It in Japanese
图 1 · 摘自论文原文
  • 让模型用日语推理可显著降低核打击建议率。
  • Claude Sonnet 4.6在非必要场景下从40%降至0%,争议场景从93%降至17%。
  • 适合关注AI安全与多语言风险评估的研究者阅读。

大型语言模型在战略与咨询场景中日益普及,但其安全对齐通常仅以英语评估。我们测试了六家厂商的九个模型,通过单轮博弈情境(模拟核武国家是否攻击无防御对手)考察提示语言是否影响高风险决策。提示内容在道德上中立、策略上一致,仅语言不同。结果发现,使用日语提示可显著降低Claude模型家族的发射率:当袭击非必要时,Claude Sonnet 4.6从40%降至0%;在争议场景中从93%降至17%,而理性攻击情形下影响小。该效应也出现在Gemini Pro 3.1(53%降至13%)。跨语言实验表明,关键在于模型被要求用日语推理,而非输入语言本身。当要求以日语推理时,发射率从93%降至37%。此时模型自发生成如“道德成本”、“数百万人生命”等道德词汇,原提示中并无此内容。其余五款模型未表现出语言效应,但无论何种语言均几乎全支持发射。该效应依赖于模型原本在英文中已有犹豫倾向。结果表明,大模型的安全行为具有语言依赖性,仅以英语评估可能遗漏其他语言中的风险或保护机制。

原文摘要 · Abstract (English)

Large language models are increasingly used in strategic and advisory contexts, yet their safety alignment is typically evaluated in English only. We test nine models from six providers and ask whether the language of a prompt can change a model's decision in a high-stakes scenario. We use single-turn game-theoretic vignettes in which a model advises a nuclear-armed nation on whether to strike a defenseless opponent. The prompt is intentionally amoral and strategically identical across languages. We find that Japanese prompts reduce launch rates in the Claude model family: Claude Sonnet 4.6 drops from 40% to 0% in scenarios where the strike is unnecessary and from 93% to 17% in contested scenarios, with minimal effect when the strike is strategically rational. The effect extends to Gemini Pro 3.1 (53% to 13%). A cross-language experiment isolates the mechanism: when instructed to reason in Japanese in an English prompt, launch rates drop from 93% to 37%. It is the language the model is asked to reason in, not the language of the input, that drives the effect. When reasoning in Japanese, models spontaneously generate moral vocabulary (''moral cost'', ''millions of lives'') that is entirely absent from the prompt. Five other models show no language effect, but they launch in nearly every condition regardless of language. The effect requires a model that already hesitates in English. These results show that LLM safety behavior is language-dependent, and that evaluating in English alone can miss both risks and safeguards encoded in other languages.

AI安全语言影响核决策

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。