arXiv:2502.20432cs.AIcs.CY2025-02NeurIPS被引 23

用行为博弈论评估大模型战略推理能力,发现模型表现不只靠规模。

LLM Strategic Reasoning: Agentic Study through Behavioral Game Theory

  • 基于行为博弈论框架,分离模型推理与上下文影响。
  • GPT-o3-mini等三模型表现领先,但规模非决定性因素。
  • 提示词效果因模型而异,且存在性别/性取向等隐性偏见。

战略决策涉及交互式推理,即智能体根据他人行为调整自身选择,但现有大语言模型(LLMs)评估多聚焦于纳什均衡(NE)逼近,忽视其战略选择背后的机制。为弥补此缺口,我们提出基于行为博弈论的评估框架,分离推理能力与上下文效应。测试22个前沿LLM后发现,GPT-o3-mini、GPT-o1和DeepSeek-R1在多数博弈中表现领先,但模型规模并非决定性能的关键。在提示增强方面,思维链(CoT)提示并非普适有效,仅对特定水平的模型提升战略推理,其余则收效甚微。此外,我们研究了编码人口特征对模型的影响,发现某些设定会改变决策模式:例如GPT-4o在赋予女性特征时表现出更强的战略推理,而Gemma则对异性恋身份赋予更高推理水平,相较其他性取向,显示出内在偏见。这些发现强调需建立伦理标准与上下文对齐机制,以平衡推理能力提升与公平性。

原文摘要 · Abstract (English)

Strategic decision-making involves interactive reasoning where agents adapt their choices in response to others, yet existing evaluations of large language models (LLMs) often emphasize Nash Equilibrium (NE) approximation, overlooking the mechanisms driving their strategic choices. To bridge this gap, we introduce an evaluation framework grounded in behavioral game theory, disentangling reasoning capability from contextual effects. Testing 22 state-of-the-art LLMs, we find that GPT-o3-mini, GPT-o1, and DeepSeek-R1 dominate most games yet also demonstrate that the model scale alone does not determine performance. In terms of prompting enhancement, Chain-of-Thought (CoT) prompting is not universally effective, as it increases strategic reasoning only for models at certain levels while providing limited gains elsewhere. Additionally, we investigate the impact of encoded demographic features on the models, observing that certain assignments impact the decision-making pattern. For instance, GPT-4o shows stronger strategic reasoning with female traits than males, while Gemma assigns higher reasoning levels to heterosexual identities compared to other sexual orientations, indicating inherent biases. These findings underscore the need for ethical standards and contextual alignment to balance improved reasoning with fairness.

战略推理行为博弈论模型偏见提示工程

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。