arXiv:2605.02469cs.LGcs.AI2026-05被引 1

提出一种匹配强化学习目标的监督微调方法,提升训练效率与稳定性。

Reference-Sampled Boltzmann Projection for KL-Regularized RLVR: Target-Matched Weighted SFT, Finite One-Shot Gaps, and Policy Mirror Descent

  • 基于参考策略采样设计加权SFT目标,使模型逼近理想奖励分布。
  • 单次训练即可达到性能饱和,避免多轮迭代的资源浪费。
  • 适合追求高效训练、对政策稳定性和覆盖范围有要求的研究者。

在线强化学习中可验证奖励(RLVR)将可检查结果转化为可扩展的训练信号,但其仍需在优化路径上进行回溯生成、验证评分和参考策略评估。预先计算回溯结果的静态加权监督微调(SFT)看似能消除瓶颈,但权重仅由奖励决定并不充分:采样器与权重共同决定了被拟合的策略。本文识别出一种参考采样加权SFT目标,其诱导策略恰好等于固定参考下KL正则化的RLVR优化器。该优化器即标准玻尔兹曼目标策略,通过对参考策略按验证奖励指数倾斜得到。使加权SFT诱导策略匹配此目标,强制密度比权重;在参考采样子类中,这唯一地(除提示缩放外)简化为提示归一化玻尔兹曼权重 $\exp(r(x,y)/β)/Z(x)$。BOLT是一种玻尔兹曼目标匹配的SFT过程,作为该投影的实证估计器。有限一次解分析将精确存储支持价格 $β\log(1/π^*(S_N\mid x))$ 与分区估计、有效样本量方差、泛化、优化及近似误差分离。该分解解释了为何额外的SFT轮次无法弥补缺失的参考策略覆盖,并揭示温度-覆盖-方差前沿。当覆盖需要自适应采样时,刷新的玻尔兹曼投影变为KL策略镜面下降;有限内部求解以附加漂移形式进入精确镜面步。单次运行的Qwen实验在限定范围内提供了目标匹配权重、单次饱和、刷新采样优势和优化时间节省的证据。

原文摘要 · Abstract (English)

Online reinforcement learning with verifiable rewards (RLVR) turns checkable outcomes into a scalable training signal, but it keeps rollout generation, verifier scoring, and reference-policy evaluations on the optimization path. Static weighted supervised fine-tuning (SFT) on precomputed rollouts seems to remove this bottleneck, yet a weighted likelihood is not specified by rewards alone: its sampler and weights induce the policy being fit. This paper identifies the reference-sampled weighted-SFT objective whose induced policy equals the fixed-reference KL-regularized RLVR optimizer. The optimizer is the standard Boltzmann target policy, obtained by exponentially tilting the reference policy by verifier reward. Matching a weighted-SFT induced policy to this target forces density-ratio weights; in the reference-sampled subclass, this reduces uniquely, up to prompt scaling, to the prompt-normalized Boltzmann weight $\exp(r(x,y)/β)/Z(x)$. BOLT, a Boltzmann-Targeted SFT procedure, is the empirical estimator of this projection. The finite one-shot analysis separates the exact stored-support price $β\log(1/π^*(S_N\mid x))$ from partition estimation, effective-sample-size variance, generalization, optimization, and approximation errors. This decomposition explains why extra SFT epochs cannot repair missing reference-policy coverage and exposes the temperature--coverage--variance frontier. When coverage needs adaptive sampling, refreshed Boltzmann projections become KL policy mirror descent; finite inner solves enter as additive drift from the exact mirror step. Single-run Qwen experiments provide projection evidence for the target-matched weight, one-shot saturation, refreshed-sampler gains, and optimization-time savings, within the stated single-run scope.

强化学习监督微调策略优化奖励建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。