arXiv:2512.07462cs.MAcs.AI2025-12被引 11

用博弈论分析大模型行为,发现其合作倾向受语言和激励影响。

Understanding LLM Agent Behaviours via Game Theory: Strategy Recognition, Biases and Multi-Agent Dynamics

  • 构建双博弈环境,分离激励强度与多智能体历史影响。
  • 发现模型在末期普遍转向背叛,跨语言行为差异显著。
  • 适合关注AI治理、多智能体安全的科研与工程人员。

随着大语言模型(LLMs)越来越多地作为交互系统与人类社会中的自主决策者运行,理解其策略性行为对安全性、协调性及人工智能驱动的社会经济基础设施设计具有重要意义。评估此类行为需超越输出文本本身,揭示其决策背后的意图。本文扩展了FAIRGAME框架,通过两个互补改进系统评估LLM在重复社会困境中的行为:一个收益量化的囚徒困境,用于分离激励幅度的影响;一个动态收益与多智能体历史结合的公共品博弈。实验揭示了跨模型与语言的一致行为特征,包括对激励敏感的合作、跨语言差异以及末期趋向背叛的对齐现象。通过在标准重复博弈策略上训练监督分类模型,并应用于FAIRGAME轨迹,结果表明LLMs表现出系统性、模型与语言依赖的行为意图,语言表述的影响有时甚至等同于架构差异。这些发现为审计大模型作为策略性代理提供了统一方法论基础,并揭示了直接关乎人工智能治理、集体决策与安全多智能体系统设计的系统性合作偏差。

原文摘要 · Abstract (English)

As Large Language Models (LLMs) increasingly operate as autonomous decision-makers in interactive and multi-agent systems and human societies, understanding their strategic behaviour has profound implications for safety, coordination, and the design of AI-driven social and economic infrastructures. Assessing such behaviour requires methods that capture not only what LLMs output, but the underlying intentions that guide their decisions. In this work, we extend the FAIRGAME framework to systematically evaluate LLM behaviour in repeated social dilemmas through two complementary advances: a payoff-scaled Prisoners Dilemma isolating sensitivity to incentive magnitude, and an integrated multi-agent Public Goods Game with dynamic payoffs and multi-agent histories. These environments reveal consistent behavioural signatures across models and languages, including incentive-sensitive cooperation, cross-linguistic divergence and end-game alignment toward defection. To interpret these patterns, we train traditional supervised classification models on canonical repeated-game strategies and apply them to FAIRGAME trajectories, showing that LLMs exhibit systematic, model- and language-dependent behavioural intentions, with linguistic framing at times exerting effects as strong as architectural differences. Together, these findings provide a unified methodological foundation for auditing LLMs as strategic agents and reveal systematic cooperation biases with direct implications for AI governance, collective decision-making, and the design of safe multi-agent systems.

大模型行为博弈论多智能体合作偏差

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。