用行为特征嵌入预测大模型策略迁移,比收益结构更关键
Strategy, Not Payoffs: A Behavioural Embedding of Normal-Form Games

- 设计两种行为特征嵌入:纳什均衡熵与最优回应敏感性
- 新嵌入能准确预测未见游戏的策略能力变化,旧方法仅记游戏身份
- 适合研究大模型策略学习、博弈推理与泛化能力的学者
学习某一策略任务不仅改变直接教学内容:在某个游戏中微调可能提升或降低智能体在其他游戏中的推理能力。理解并预测这种策略能力的迁移,仍是大语言模型(LLMs)的核心挑战。标准型博弈因其明确的收益设定和可分析的均衡行为,成为研究该现象的理想实验场。本文探究游戏嵌入是否能解释并预测大模型在不同游戏间微调后的策略能力变化。我们提出一种轻量级双特征嵌入,捕捉核心行为需求:纳什均衡的熵与最优回应对对手动作的敏感性。结果显示,现有公开结构嵌入主要记忆游戏身份,无法泛化;而我们的行为嵌入可可靠预测未见游戏的表现变化。这表明大模型策略能力的迁移并非由博弈收益几何决定,而是由其要求的决策行为结构所主导。
原文摘要 · Abstract (English)
Learning a strategic task changes more than what is directly taught: fine-tuning on one game can either enhance or degrade an agent's ability to reason in another. Understanding and predicting this transfer of strategic capabilities, however, remains a key challenge for large language models (LLMs). Normal-form games provide an ideal testbed for analyzing this phenomenon, as they feature explicitly defined payoffs and well-characterized equilibrium behaviours. In this work, we investigate whether game embeddings can explain and predict changes in LLM strategic capabilities following fine-tuning across different games. We propose a lightweight two-feature embedding that captures fundamental behavioural demands: the entropy of the Nash equilibrium and the sensitivity of optimal responses to an opponent's action. We show that while existing published structural embeddings primarily memorize game identities and fail to generalize, our behavioural embedding reliably predicts performance changes on held-out games. These results demonstrate that the transfer of strategic capabilities in LLMs is not dictated by the payoff geometry of a game, but by the underlying structure of the decision-making behaviour it requires.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。