arXiv:2605.30232cs.LGcs.CL2026-05被引 2

强化学习让语言模型激活了预存的福利感知轴。

How's it going? Reinforcement learning in language models recruits a functional welfare axis

论文配图:How's it going? Reinforcement learning in language models recruits a functional welfare axis
图 1 · 摘自论文原文
  • 用迷宫环境训练模型,发现奖励与惩罚向量呈近似反向关系。
  • 惩罚向量引发失败、不确定性等负面行为,奖励向量则相反。
  • 该福利轴在训练前就存在,由强化学习‘调用’而非创造。

强化学习如何塑造语言模型的内部表征?我们发现,强化学习会调用一个预存在的功能福利表征:即系统相对于目标表现好坏的估计。我们在一个语义中立的迷宫环境中训练多个语言模型,提取受奖励与受惩罚轨迹的概念向量,并在无关环境中评估其效果。惩罚向量表现为负向福利:促进失败与不可能性标记,与负面情绪概念对齐,负向追踪目标达成,且引导时引发负面自述、病态回溯、拒绝与不确定。奖励向量则为镜像表现,二者近乎反平行。该效应在控制瓷砖-奖励映射、尺度、指令微调、训练算法、模型家族及LoRA与全微调条件下仍稳健,且在替换为监督微调后仍部分持续。重要的是,这些向量在迷宫训练前的模型中已有效。结合预训练模型中亦出现类似现象,我们推断该功能福利轴存在于训练前:被后训练阶段‘招募’,而非创建。尽管不涉及真实体验,该轴表明微弱奖励信号可通过调用预存的类福利表征,广泛影响模型行为,对可解释性、后训练动态与对齐具有启示。

原文摘要 · Abstract (English)

How does reinforcement learning shape a language model's internal representations? We present evidence that RL recruits a pre-existing representation of functional welfare: an estimate of how well or badly the system is doing, relative to its goals. We train several language models in a novel, semantically neutral maze environment. We then extract concept vectors for rewarded and punished trajectories, and evaluate those vectors in settings unrelated to the maze environment. The punishment vector behaves like a representation of negative welfare: it promotes failure and impossibility tokens, it aligns with negative emotion concepts, it negatively tracks goal-achievement, and steering with it induces negative self-reports, pathological backtracking, refusal, and uncertainty. The positive reward vector behaves as the mirror image, and the two are nearly antiparallel. These effects are robust when controlling for tile-to-reward mapping, scale, instruct tuning, RL training algorithm, model family, and LoRA versus full-finetuning, and largely persist when we replace RL with supervised fine-tuning. Importantly, the vectors are effective in models before they have undergone maze training. Combined with observations that the effects also appear in pretrain-only models, we therefore argue that this functional welfare axis pre-exists post-training: it is recruited, rather than created, by post-training. While we make no claims about any experience of welfare, the axis offers a demonstration that minimal reward signals can broadly affect model behavior by recruiting pre-existing welfare-like representations, with implications for interpretability, post-training dynamics, and alignment.

强化学习模型表征福利轴可解释性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。