语言模型在无语义提示下难以进行博弈推理,残差信号可揭示三种策略模式。
Equilibrium Residuals Expose Three Regimes of Matrix-Game Strategic Reasoning in Language Models
- 用生成的零和矩阵游戏测试模型,发现去语义后准确率骤降
- 通过残差奖励训练,5×5以上博弈成功率从2%提升至61%
- 残差信号对收益扰动稳定,适合跨规模迁移,但受输出格式限制
大型语言模型在包含语义线索的博弈基准上表现良好,但移除语义后战略计算能力显著下降。我们通过程序生成的零和矩阵游戏验证此现象:模型在匿名的2×2、3×3、5×5支付矩阵上的成功率分别降至34%、18%和2%。该基准分离出语义回忆、近似纳什均衡计算与输出接口瓶颈三类因素。仅在2×2和3×3游戏上训练,监督微调使未见的5×5至7×7博弈成功率从2%提升至61%,而基于可被利用性奖励的训练平均达37%且种子间方差高。我们证明,可被利用性残差在收益扰动下是2-利普希茨连续的,不同于不连续的顶点返回线性规划均衡选择器,解释了为何残差训练能在支付变化下实现迁移,尽管格式不稳定限制了均值性能。通过主导动作填充实验,我们提供因果证据:训练模型能解决嵌入大矩阵中的3×3博弈,而随机填充对照组失败,密集的12×12博弈仍接近失败。因此,程序化评估对于衡量战略推理至关重要,残差奖励揭示了一条真实但受限于格式的近似均衡计算路径。
原文摘要 · Abstract (English)
Large language models can score well on named game-theory benchmarks while failing on the same strategic computation once semantic cues are removed. We show this gap with procedurally generated zero-sum matrix games: a model that recognizes familiar games drops to 34%, 18%, and 2% success on anonymous $2{\times}2$, $3{\times}3$, and $5{\times}5$ payoff matrices. The benchmark separates semantic recall, learned approximate Nash computation, and an output-interface bottleneck that limits scale. Training only on $2{\times}2$ and $3{\times}3$ games, supervised fine-tuning raises unseen $5{\times}5$--$7{\times}7$ success from 2% to 61%, while exploitability-reward training averages 37% with high seed variance. We prove that the exploitability residual is $2$-Lipschitz in payoff perturbations, unlike discontinuous vertex-returning LP equilibrium selectors, explaining why residual training can transfer under payoff shifts even when formatting instability limits mean performance. A dominated-action padding experiment provides causal evidence: trained models solve $3{\times}3$ games embedded in much larger matrices, while random-padded controls fail and dense $12{\times}12$ games remain near failure. Procedural evaluation is therefore necessary for measuring strategic reasoning, and residual rewards expose a real but format-limited route to approximate equilibrium computation.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。