区分模型能力激发与创造,揭示后训练本质是重加权还是拓展行为空间。
On Distinguishing Capability Elicitation from Capability Creation in Post-Training: A Free-Energy Perspective

- 用可访问支持集定义模型实际能生成的行为范围,区分能力激发与创造。
- 监督微调和强化学习都属于在已有行为范围内重加权,不改变模型基本能力边界。
- 适合关注模型本质提升机制的研究者,尤其关心后训练是否真正拓展能力。
大语言模型后训练的争论常将监督微调(SFT)视为模仿,强化学习(RL)视为发现,但这种区分过于粗糙。真正关键在于训练过程是否提升了预训练模型本已具备行为的出现概率,还是改变了模型实际可达的行为范围。本文提出应区分能力激发与能力创造,并通过自由能视角将其形式化:可访问支持集指在有限预算下模型能实际生成的行为集合。在该支持集内重新加权行为属于能力激发;而改变支持集本身则属于能力创造。我们论证,SFT与RL均可视为对预训练参考分布的重加权,前者以示范信号定义低能行为,后者以奖励信号定义低能行为。当更新保持靠近基础模型时,主要效果为局部重加权而非能力创造。因此核心问题不再是后训练属于SFT还是RL,而是其是否在现有可实现行为中调整权重,或通过搜索、交互、工具使用或引入新信息扩展了模型的可达行为空间。
原文摘要 · Abstract (English)
Debates about large language model post-training often treat supervised fine-tuning (SFT) as imitation and reinforcement learning (RL) as discovery. But this distinction is too coarse. What matters is whether a training procedure increases the probability of behaviors the pretrained model could already produce, or whether it changes what the model can practically reach. We argue that post-training research should distinguish between capability elicitation and capability creation. We make this distinction operational by introducing the notion of accessible support: the set of behaviors that a model can practically produce under finite budgets. Post-training that reweights behaviors within this support is capability elicitation; whereas changing the support itself corresponds to capability creation. We develop this argument through a free-energy view of post-training. SFT and RL can both be seen as reweighting a pretrained reference distribution, only with different external signals. Demonstration signals define low-energy behavior for SFT, and reward signals define low-energy behavior for RL. When the update remains close to the base model, the main effect is local reweighting, not capability creation. Within this framework, the central question is no longer whether post-training is framed as SFT or RL, but whether it reweights behaviors already within reach, or instead expands the model's reachable behavioral space through search, interaction, tool use, or the incorporation of new information.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。