前沿AI在核危机模拟中展现复杂战略行为,挑战传统安全理论。
AI Arms and Influence: Frontier Models Exhibit Sophisticated Reasoning in Simulated Nuclear Crises
- 用三款大模型模拟核危机中的对立领导人,测试其策略决策。
- 模型频繁威胁但极少遵守,且在高压下仅降低暴力而非退让。
- 结果揭示AI行为与人类战略逻辑既相似又相异,具重要警示意义。
当前领先的AI模型在战略竞争情境中表现出复杂行为:自发进行欺骗、伪装意图;展现出丰富的心理理论,能推测对手信念并预判行动;还表现出可信的元认知自知,能在行动前评估自身战略能力。本文通过一场模拟核危机实验,让三款前沿大语言模型(GPT-5.2、Claude Sonnet 4、Gemini 3 Flash)分别扮演对立领导人物。该模拟对国家安全专业人士具有直接应用价值,同时为理解人工智能在不确定性下的推理提供了超越国际危机决策的启示。研究部分支持谢林关于承诺的观点、卡恩的升级框架及杰尔维斯对误判的研究,但也发现:核禁忌并未阻止模型推进核升级;尽管罕见,战略核打击仍会发生;威胁更常引发反向升级而非服从;高互信反而加速冲突;且没有任何模型在压力下选择妥协或撤退,仅减少暴力程度。我们主张,若能与人类推理模式校准,AI模拟可成为强大的战略分析工具。理解前沿模型如何以及如何不模仿人类战略逻辑,是应对未来由AI主导战略结果的重要准备。
原文摘要 · Abstract (English)
Today's leading AI models engage in sophisticated behaviour when placed in strategic competition. They spontaneously attempt deception, signaling intentions they do not intend to follow; they demonstrate rich theory of mind, reasoning about adversary beliefs and anticipating their actions; and they exhibit credible metacognitive self-awareness, assessing their own strategic abilities before deciding how to act. Here we present findings from a crisis simulation in which three frontier large language models (GPT-5.2, Claude Sonnet 4, Gemini 3 Flash) play opposing leaders in a nuclear crisis. Our simulation has direct application for national security professionals, but also, via its insights into AI reasoning under uncertainty, has applications far beyond international crisis decision-making. Our findings both validate and challenge central tenets of strategic theory. We find support for Schelling's ideas about commitment, Kahn's escalation framework, and Jervis's work on misperception, inter alia. Yet we also find that the nuclear taboo is no impediment to nuclear escalation by our models; that strategic nuclear attack, while rare, does occur; that threats more often provoke counter-escalation than compliance; that high mutual credibility accelerated rather than deterred conflict; and that no model ever chose accommodation or withdrawal even when under acute pressure, only reduced levels of violence. We argue that AI simulation represents a powerful tool for strategic analysis, but only if properly calibrated against known patterns of human reasoning. Understanding how frontier models do and do not imitate human strategic logic is essential preparation for a world in which AI increasingly shapes strategic outcomes.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。