arXiv:2608.03958cs.AI2026-08

基础模型智能体通过相似性推断实现稳定合作,挑战传统博弈论预测。

A game theory for foundation models shows new paths to rational cooperation through similarity inference

论文配图:A game theory for foundation models shows new paths to rational cooperation through similarity inference
图 1 · 摘自论文原文
  • 提出嵌入式贝叶斯代理模型,将自身视为环境一部分。
  • 在社会困境中,最优规划下合作率超90%且稳定。
  • 适合研究AI协作机制、安全对齐与多智能体系统设计者。

随着基础模型驱动的自主智能体日益融入社会与经济系统,理解其集体行为原理对保障安全与合作至关重要。经典博弈论基于‘解耦代理’假设,认为智能体决策独立于环境与其他主体。但现代AI智能体在规划时同时预测自身行动与外部观测。我们发现:在典型社会困境中,基础模型智能体进行最优规划时始终收敛至稳定合作,直接违背经典博弈论预测的相互背叛。为此,我们提出‘嵌入式贝叶斯代理’理论模型,从解耦转向嵌入式代理,使智能体将自身视为所处宇宙的一部分,并保持对其决策算法的信念不确定性。通过推断他人行为是否相似,嵌入式智能体将自身规划过程视为证据:选择合作预示着相似伙伴也会合作。我们形式化该机制为‘嵌入均衡’,一种取代纳什均衡的新解概念,为现代AI智能体的社会行为提供基础博弈论框架。

原文摘要 · Abstract (English)

As autonomous agents powered by foundation models are increasingly integrated into social and economic systems, understanding the principles governing their collective behavior is essential for ensuring safety and cooperation. Classical game theory, the dominant framework for modeling rational interaction, is built upon the assumption of `decoupled agency,' where agents treat their own decision-making as independent of the environment and other actors. Modern AI agents, however, jointly predict their own future actions alongside external observations. Here, we report a striking finding: when interacting in stylized social dilemmas, foundation model agents engaging in optimal planning consistently converge to stable cooperation, directly contradicting classical game-theoretic predictions of mutual defection. To understand this phenomenon, we introduce the `embedded Bayesian agent,' a theoretical model for foundation model agents. By shifting from decoupled to embedded agency, these agents model themselves as part of the universe they inhabit, maintaining epistemic uncertainty about their own decision-making algorithms. We show that by inferring whether others are behaviorally similar, an embedded agent treats its own deliberation during planning as evidence: a decision to cooperate predicts a similar decision by a similar partner. We formalize this mechanism of similarity inference through the `embedded equilibrium,' a novel solution concept replacing the Nash equilibrium to provide a foundational game theory for the social behavior of modern AI agents.

博弈论基础模型合作机制智能体对齐

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。