arXiv:2602.10885cs.AIcs.LG2026-02被引 17

用自演化评分标准自动奖励思维链,无需人工标注。

Reinforcing Chain-of-Thought Reasoning with Self-Evolving Rubrics

  • 自动生成并动态更新评分标准,替代人工标注。
  • 在无结果奖励时仍能有效提升思维链质量。
  • 可作为提示词提升推理性能,适合大模型优化研究者。

尽管思维链(CoT)在大语言模型推理中至关重要,但直接对其奖励却面临挑战:训练奖励模型需要大量人工标注,而静态奖励模型难以适应思维链分布的变化并易受奖励黑客攻击。为此,我们提出一种无需人工标注且可自主演化的思维链奖励方法——RLCER(基于自演化评分标准的强化学习)。该方法通过自生成、自演化的评分标准对思维链进行奖励,增强了以结果为中心的强化学习与价值回归(RLVR)的效果。实验表明,即使在没有结果奖励的情况下,自演化评分标准仍能提供可靠的思维链监督信号,使RLCER优于传统结果导向的RLVR。此外,这些自生成的评分标准作为提示使用时,还能进一步提升推理阶段的表现。

原文摘要 · Abstract (English)

Despite chain-of-thought (CoT) playing crucial roles in LLM reasoning, directly rewarding it is difficult: training a reward model demands heavy human labeling efforts, and static RMs struggle with evolving CoT distributions and reward hacking. These challenges motivate us to seek an autonomous CoT rewarding approach that requires no human annotation efforts and can evolve gradually. Inspired by recent self-evolving training methods, we propose \textbf{RLCER} (\textbf{R}einforcement \textbf{L}earning with \textbf{C}oT Supervision via Self-\textbf{E}volving \textbf{R}ubrics), which enhances the outcome-centric RLVR by rewarding CoTs with self-proposed and self-evolving rubrics. We show that self-proposed and self-evolving rubrics provide reliable CoT supervision signals even without outcome rewards, enabling RLCER to outperform outcome-centric RLVR. Moreover, when used as in-prompt hints, these self-proposed rubrics further improve inference-time performance.

思维链强化学习自演化大模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。