揭示大模型决策背后的内在机制差异
The Internal Anatomy of Strategic Choice in Large Language Models
- 通过追踪激活值分析模型在博弈中的战略选择过程
- 基础与指令微调模型行为相似,但激励传递路径不同
- 训练方式改变决策路径,但外表行为不变
大型语言模型作为策略性代理和人类选择的模型,其决策行为虽类似策略代理,但计算方式不同。我们记录了四种开源模型(密集型与混合专家架构,含一对基线-指令微调模型)在144个严格序数2×2博弈中的一次性决策过程中的激活值。从提示中的激励信号,经由激活状态,到最终选择,全程追踪。所有模型均能检测到激励信号与最终选择,但激励是否传达到选择、是否对齐,以及增强激励是否会改变偏好,各模型表现不一。基线与指令微调的Qwen2.5模型在初始状态下几乎完全一致,但在激励能否到达决策环节上存在差异。固定决策线索在内部可区分,但仅选择性地影响行为。相同行为可能基于不同的内部计算;后训练可重塑从激励表征到决策的路径,而行为和可解码信息基本保持不变。
原文摘要 · Abstract (English)
Large language models act as strategic agents and models of human choice, yet choosing like a strategic agent does not mean computing like one. We recorded activations from four open-weight models --- dense and mixture-of-experts, including a matched base--instruct pair --- in one-shot play of 144 strict ordinal $2\times2$ games. We followed a prespecified incentive from prompt, through activations, to choice. Dense models mirrored the unadjusted human decline with game complexity. Incentive and choice were detectable in every model, but models differed in whether incentive reached the choice, aligned with it and, where tested, whether strengthening it shifted preference. The base and instruction-tuned Qwen2.5 models chose almost identically at baseline yet differed in whether incentive reached choice. Fixed decision cues were distinguishable internally but changed choices selectively. Similar behaviour can rest on different computation; post-training can reshape the path from represented incentive to decision while leaving behaviour and decodable information largely intact.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。