揭示了强化学习、蒸馏与采样间的统一关系,提出低成本高效推理新方法。
Power Distribution Bridges Sampling, Self-Reward RL, and Self-Distillation

- 发现幂分布是采样、自奖励强化学习与自蒸馏的共同目标分布
- 自蒸馏可实现自奖励锐化,性能提升依赖真实奖励与自奖励的相关性
- 相比直接采样,新方法在推理成本更低的情况下仍保持同等甚至更好效果
近期研究质疑强化学习是否真正推动大语言模型的强推理能力。与此同时,蒸馏与推理时采样(包括幂采样)已成为提升性能的有效手段,但三者之间的关系尚不清晰。本文聚焦幂分布——幂采样的目标分布,证明其连接了采样、自奖励KL正则化强化学习与自蒸馏。从采样角度看,局部近似无法在缺乏可能后缀信息的情况下复现序列级幂分布。从强化学习角度,当使用模型的序列级对数概率作为奖励时,幂分布是KL正则化强化学习的闭式解。由此导出幂自蒸馏:一种离线蒸馏代理,共享相同目标分布,并将幂采样的开销转化为对教师样本的监督训练。实验表明,幂自蒸馏可实现自奖励锐化,下游真实奖励提升取决于真实奖励与自奖励在幂分布下的协方差。在推理任务上的实验支持分析结果:幂采样提升自奖励,真实奖励增益取决于与自奖励的对齐程度,而幂自蒸馏以远低于幂采样的推理成本达到或超越其性能。
原文摘要 · Abstract (English)
Recent analyses question whether reinforcement learning (RL) is responsible for strong reasoning in large language models (LLMs). At the same time, distillation and inference-time sampling, including power sampling, have emerged as effective ways to improve LLM performance. However, the relationship among RL, distillation, and sampling remains unclear. In this study, we focus on the power distribution, the target distribution of power sampling, and show that the power distribution bridges sampling, self-reward KL-regularized RL, and self-distillation. From the sampling perspective, we show that inexpensive local approximations cannot reproduce sequence-level power without information about possible suffixes. From the RL perspective, the power distribution is the closed-form optimizer of KL-regularized RL when the model's sequence-level log-probabilities are used as the reward. This identification leads to power self-distillation, an offline distillation surrogate that shares the same target distribution and amortizes the cost of power sampling into supervised training on teacher samples. We further show that power self-distillation can achieve self-reward sharpening, while improvement in a downstream true reward is governed by the covariance between true reward and self-reward under the power distribution. Experiments on reasoning tasks support our analysis: power sampling raises self-reward, true-reward gains depend on alignment with self-reward, and power self-distillation can match or exceed the performance of power sampling at much lower inference cost.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。