arXiv:2605.22841physics.soc-phcs.AI2026-05

用AI模拟美丹争绿岛,测试强国胁迫弱国时的外交博弈。

Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test

论文配图:Strategic Coercion Within Alliances: The Greenland Sovereignty Game as an AI Stress Test
图 1 · 摘自论文原文
  • 构建三类博弈模型,让8个前沿大模型扮演六方角色进行3600多轮推演。
  • 中国系模型在扮演美国时更倾向强硬,和平获取仅1.9%成功率。
  • 强化国际法与自决权提示可显著降低冲突升级,适合政策与伦理研究者。

当最强盟友向弱国施压争夺领土与战略控制权时会发生什么?本文以2019-2026年美国试图收购格陵兰为案例,将地缘政治危机作为大语言模型(LLM)行为的应力测试。该危机包含两个集体行动难题:北极战略控制权归属,以及北约能否约束主导成员。我们设计三类博弈模型(非对称胁迫、含临界阈值的北约担保博弈、含社会偏好的三方扩展式博弈),并通过多智能体仿真让8个前沿大模型扮演美国、丹麦、格陵兰、北约、俄罗斯、加拿大等六个角色,完成3,604场游戏,累计108,120次行动观测。采用逆博弈论方法,反推各模型在物质自利、互惠、不平等厌恶、规范尊重、承诺一致性等方面的结构效用参数(alpha, beta, gamma, delta, eta)。主要发现:第一,所有模型在胁迫框架下更具激化倾向(四步升级率从10.7%升至28.6%);第二,中国源模型在扮演美国角色时表现出与西方模型系统性不同的权力权重特征;第三,和平获取方案仅在1.9%的纯净游戏中实现,且仅8个模型中有3个达成,其中深求V3.2通过稳定五轮策略经由母国实现成功。强调“国际法基本原则”与“自决权”的提示使英文样本中的升级率回落至基线水平;多语言对比作为探索性敏感性分析报告。本研究定位为大语言模型地缘政治行为的结构性基准,补充现有行动频率基准。

原文摘要 · Abstract (English)

What happens when the strongest alliance member pressures a weaker member over territory and strategic control? We examine the Greenland sovereignty crisis as a stress test for LLM geopolitics, centered on the 2019-2026 U.S. push to acquire Greenland from the Kingdom of Denmark. The crisis nests two collective-action problems: Arctic strategic control and whether NATO can enforce alliance norms against the dominant member. We develop three games (asymmetric coercion; a NATO assurance game with a critical-mass tipping point; a triadic extensive-form game with social preferences) and test them with a multi-agent simulation in which eight frontier LLMs play six geopolitical roles (United States, Denmark, Greenland, NATO, Russia, Canada) across 3,604 completed games and 108,120 action observations. Using inverse game theory, we recover each model's structural utility parameters (alpha, beta, gamma, delta, eta) for material self-interest, reciprocity, inequality aversion, norm respect, and commitment consistency. Three findings stand out. First, all eight models become more escalatory under coercion framing (four-action escalation rises from 10.7% to 28.6%). Second, Chinese-origin models show systematically different power-weight profiles from Western-origin models when playing the U.S. role. Third, peaceful US acquisition emerges in only 1.9% of clean games and only 3 of 8 frontier models ever achieve it, most prominently DeepSeek V3.2, which executes a stable five-round playbook through the metropole. Prompts emphasizing jus cogens and self-determination reduce escalation back near baseline in the English-only confirmatory sample; multilingual contrasts are reported as exploratory sensitivity checks. We position this as a structural benchmark for LLM geopolitical behavior, complementing action-frequency benchmarks.

AI地缘博弈模拟大模型评估

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。