arXiv:2608.01193cs.AIcs.CY2026-08

测试大模型在模拟AI竞赛中的安全策略,发现表现差异大且不可靠。

Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races

论文配图:Humans Are More Diverse: Frontier LLMs Show Extreme Policies in Idealised AI Development Races
图 1 · 摘自论文原文
  • 用博弈论框架测试多智能体的决策行为,验证规则理解与状态追踪能力。
  • 七种模型中行动序列差异显著,即使规则相同,结果也因模型而异。
  • 强调需做有效性检查,避免误判模型具备人类级战略思维。

AI开发竞赛形成多智能体安全困境:每家公司可选择缓慢安全发展,或加速冒险以获取潜在奖励。我们通过重复博弈研究了两到五名参与者在该情境下大语言模型(LLM)代理的行为策略。但有效动作不等于真正理解游戏规则,因此我们在行为分析前设置审计关卡:先验证游戏引擎,再测试规则记忆、状态追踪、收益计算及不同等效任务描述下的稳定性。随后将模型行动序列与演化博弈理论基准及已发表的人类数据对比,分析模型、风险条件、角色设定及两至五人竞赛中的差异。审计发现,强规则记忆可与弱状态追踪和预期收益计算并存;经验证的算术处理与响应表示方式改变,即便规则不变,后续行为亦受影响。在七个测试模型端点中,总体率掩盖了行动序列、对对手反应及对竞赛位置响应的巨大差异。三至五人竞赛模式的表现亦为模型特异性,非简单增加竞争者的单一效应。这些结果表明,多智能体AI竞赛模拟必须进行有效性检验和轨迹级分析,才能可靠描述其输出为战略性、类人或安全意识行为。研究结果为探索性,仅适用于所测试模型、提示和解码设置。

原文摘要 · Abstract (English)

An AI development race creates a multi-agent safety dilemma. Each company can develop slowly and safely, or move faster while taking a risk that may remove its final reward. We use this repeated game to study strategic safety behaviour among large language model (LLM) agents in races with two to five players. However, a valid action does not show that an agent understands the game. We therefore place an audit gate before behavioural interpretation. We first verify the game engine, then test rule recall, state tracking, payoff calculation, and stability under different but equivalent task descriptions. We then compare LLM action sequences with an evolutionary game-theory benchmark and published human data, and explore differences across models, risk conditions, personas, and two- to five-player races. The audit shows that strong rule recall can coexist with weak state tracking and expected-payoff calculation. Providing verified arithmetic and changing the response representation can also change later actions, even when the game rules stay fixed. Across seven tested model endpoints, aggregate rates hide large differences in action sequences, responses to opponents, and responses to race position. Patterns across the tested three- to five-player races are also model-specific rather than a single effect of adding competitors. These results show why multi-agent AI-race simulations need validity checks and trajectory-level analysis before their outputs are described as strategic, human-like, or safety-aware. Our findings are exploratory and apply only to the tested models, prompts, and decoding settings.

多智能体博弈论安全评估大模型行为

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。