arXiv:2605.27701cs.AI2026-05

通过利用奖励梯度提升大模型生成高分内容,训练更快更有效。

Cross-Entropy Games and Frost Training

  • 基于嵌入空间的奖励梯度优化策略
  • 在best-of-k设置中实现更高得分和更快收敛
  • 适用于需要高质量输出的LLM评估任务

我们提出Frost Training,一种用于改进蒙特卡洛策略优化的方法,适用于一类称为交叉熵博弈的大型语言模型作为裁判任务。核心思想是利用奖励函数在嵌入空间中的梯度信号。该信号曾被用于贪婪坐标梯度(GCG)越狱技术,我们首次证明其也可用于提升模型训练效果。通过GRPO训练进行最大似然填充验证,Frost Training显著提升了模型生成高分输出的能力,在best-of-k设置下达到更高最大分数,且训练速度更快。

原文摘要 · Abstract (English)

We present Frost Training, a method for improving Monte Carlo-based policy optimization for a large family of LLM-as-a-judge tasks called Cross-Entropy Games. The key idea is to exploit the gradient of the reward function in embedding space. This signal is used in the Greedy Coordinate Gradient (GCG) jailbreaking technique; we demonstrate for the first time that it can also be used to boost model training. We validate our method using GRPO training for maximum-likelihood infilling. Frost Training improves the model's ability to generate high-scoring outputs, reaching higher maximum scores in a best-of-k setting, and does so at an increased speed.

强化学习大模型训练奖励建模生成优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。