arXiv:2510.10813cs.AIcs.GT2025-10被引 2

大模型看似会策略思考,实则脆弱易崩,复杂时会乱用套路。

The Fragility of Strategic Thinking in Large Language Models

  • 拆解推理三步:猜对手、评估选择、最优回应,测试策略能力
  • 简单场景能猜对手并最优回应,但复杂时转向随机套路
  • 其思维偏差与人类不同,不宜直接当战略伙伴用

大型语言模型(LLMs)越来越多地应用于需要推理其他智能体行为的领域,如谈判、政策设计和市场模拟。然而,我们能否信任它们在复杂情境中进行策略思考?现有研究多关注模型是否遵循均衡策略或推理深度,却未检验其是否具备形成一致假设、基于假设评估行动并做出最优回应的能力。本文通过一系列非合作环境,构建框架分离信念形成、评估与决策过程,在静态完全信息博弈中识别该能力。结合模型输出的选择与推理轨迹,并引入新上下文无关游戏以排除记忆模仿,结果表明:当前前沿大模型虽具备真实但脆弱的策略思维——在无约束下能根据对手行为形成条件性假设并做出最优回应;但在复杂度上升时,显式递归被模型特异性逻辑转换与启发式规则取代,且这些规则既不等同于均衡推理,也不同于人类典型认知偏差。此现象已在无噪声环境下显现,提示在复杂系统中使用大模型作为战略代理需谨慎。

原文摘要 · Abstract (English)

Large Language Models (LLMs) are increasingly applied to domains that require reasoning about other agents' behavior, such as negotiation, policy design, and market simulation. However, can we trust LLMs to think strategically in complex situations? Existing research mostly evaluates LLMs' adherence to equilibrium play or their exhibited depth of reasoning, leaving open whether they display strategic thinking meant as the ability to form coherent conjectures about other agents, to evaluate possible actions conditional on those conjectures, and to best respond to them. We develop a framework to identify this ability by disentangling belief formation, evaluation, and choice in static complete-information games across a series of non-cooperative environments. By jointly analyzing models' revealed choices and reasoning traces, and introducing a new context-free game to rule out imitation from memorization, we show that strategic thinking in current frontier LLMs is real but fragile: models execute best responses to exogenous conjectures and form opponent-contingent conjectures when left unconstrained. Yet under increasing complexity explicit recursion gives way to model-specific logic shifts and heuristic rules of choice, both within and outside equilibrium reasoning. Further, these heuristics do not map directly onto the systematic biases typically observed in human strategic behavior. These findings, already emerging in noiseless settings, warrant caution in the application of LLMs as strategic agents in complex environments.

策略推理大模型缺陷博弈论AI可信

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。