用能量模型理论解释强化学习微调语言模型的内在机制
A Theoretical Lens for RL-Tuned Language Models via Energy-Based Models
- 基于能量模型构建统一变分分析框架,揭示最优策略结构
- 证明指令微调模型收敛至高质量稳定分布,且混合速度由谱间隙决定
- 解释推理模型中的熵-准确率权衡,适用于研究强化学习优化机制的学者
通过KL正则化强化学习训练的大规模语言模型展现出优秀的指令遵循、自我修正和推理能力,但其理论基础仍不充分。本文利用最优KL正则化策略的闭式能量模型(EBM)结构,对语言模型提供统一的变分分析。在奖励势函数与预训练对称性的自然假设下,我们证明指令微调模型的转移核满足关于响应质量标量势的能量平衡,从而实现单调的KL收敛至高质量稳态分布,到达优质状态的命中时间有界,混合速率由谱间隙决定。对于采用可验证奖励训练的推理模型(RLVR),目标等价于向最优推理分布最小化期望KL,次优性差距退化为自然梯度流中目标与当前准确率之间的伯努利KL。这有助于解释经验上的熵-准确率权衡现象。
原文摘要 · Abstract (English)
Large language models (LLMs) trained via KL-regularized reinforcement learning demonstrate strong instruction following, self-correction, and reasoning abilities. Yet their theoretical underpinnings remain limited. We exploit the closed-form energy-based model (EBM) structure of the optimal KL-regularized policy to provide a unified variational analysis of LLMs. For instruction-tuned models, under natural assumptions on reward potentials and pretraining symmetry, we prove that the transition kernel satisfies detailed balance with respect to a scalar potential encoding response quality. This yields monotonic KL convergence to a high-quality stationary distribution, bounded hitting times to superior states, and exponential mixing governed by the spectral gap. For reasoning models trained with verifiable rewards (RLVR), we show the objective is equivalent to expected KL minimization toward an optimal reasoning distribution, with the suboptimality gap reducing to the Bernoulli KL between target and current accuracies along the natural gradient flow. This helps explain empirical entropy-accuracy trade-offs.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。