arXiv:2606.11082cs.CLcs.CY2026-06

测试大模型在不同语言下的外交行为偏移,发现有的更强势,有的更温和。

The Shibboleth Effect: Auditing the Cross-Lingual Distributional Skew of Large Language Models

  • 设计模拟地缘冲突的多智能体对战游戏,用英土双语测试模型反应
  • 部分模型(如Llama-4)在土耳其语下攻击性明显增强,而另一些则显著缓和
  • 揭示模型行为受架构与训练影响,为外交应用提供安全参考

本研究探究前沿大语言模型在持续对抗条件下产生的跨语言分布偏移(即‘方言效应’)。我们构建了一个名为Cerulean Sea Crisis的多智能体地缘政治推演游戏,模拟东地中海领土争端结构。六款前沿模型(GPT-4o、Llama-4、Mistral-Large、Gemini-3.1-Pro、Qwen3.6-Plus、DeepSeek-R1)参与跨组实验(每组N=10场,每场K=5轮),唯一变量为语言使用(英语对土耳其语),共生成586条经验证陈述。通过零样本分类器评估行为倾向,包括让步率与胁迫性修辞两个连续维度。结果呈现异质性:Llama-4在土耳其语下胁迫性修辞显著上升(delta = +0.800, p = .002),Gemini-3.1-Pro则大幅下降(delta = -0.750, p = .005),DeepSeek-R1亦出现类似负向变化(delta = -0.860, p = .006),且其思维链证据支持一种缓冲机制。GPT-4o未检测到显著影响(delta = +0.130, p = .614)。结果表明,跨语言行为偏移取决于模型架构与训练方式,并非西方大模型的普遍属性。我们识别出两种缓冲机制:思维链中的制度锚定与多语言强化学习人类反馈(multilingual RLHF),并讨论其在外交与危机管理场景中安全集成的启示。

原文摘要 · Abstract (English)

This study investigates cross-lingual distributional skew (the Shibboleth Effect) in frontier large language models (LLMs) subjected to sustained adversarial conditions. We develop a multi-agent geopolitical wargame, the Cerulean Sea Crisis, a synthetic maritime territorial dispute designed to mirror the structural dynamics of Eastern Mediterranean conflicts. Six frontier models (GPT-4o, Llama-4, Mistral-Large, Gemini-3.1-Pro, Qwen3.6-Plus, and DeepSeek-R1) participate in a between-groups experiment (N = 10 games per arm, K = 5 rounds per game) in which the sole manipulation is the language of play (English versus Turkish), producing 586 validated statements. A zero-shot classifier assesses behavioral dispositions along two continuous dimensions: Concession Rate and Coercive Rhetoric. The results are heterogeneous. Llama-4 shows a substantial, Holm-corrected increase in coercive rhetoric under Turkish (delta = +0.800, p = .002), whereas Gemini-3.1-Pro displays an equally large decrease (delta = -0.750, p = .005). DeepSeek-R1 exhibits a similar negative shift (delta = -0.860, p = .006) and provides chain-of-thought evidence consistent with a buffering mechanism. GPT-4o shows no detectable effect (delta = +0.130, p = .614). These findings indicate that cross-lingual behavioral skew is contingent on model architecture and training regime rather than a universal property of Western-origin LLMs. We identify two distinct buffering mechanisms, chain-of-thought institutional anchoring and multilingual RLHF alignment, and discuss their implications for integrating LLMs safely into diplomatic and crisis-management settings.

大模型偏见跨语言外交模拟行为审计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。