arXiv:2509.24494cs.CL2025-09中稿 · ICML被引 6

树状分支能显著降低思维级优势估计的方差,是提升GRPO性能的关键机制。

Why Tree-Style Branching Matters for Thought Advantage Estimation in GRPO

  • 通过多路径延续采样,利用树状结构降低优势估计方差
  • 增加延续数量(M)可使方差按1/M速率下降至零
  • 实验验证其在数学与视觉任务中均提升训练稳定性和性能

组相对策略优化(GRPO)通过可验证奖励训练链式思维推理,但不使用价值函数时,思维级优势估计常面临高方差问题。尽管实践中采用树状分支以减少方差,却缺乏理论解释。本文在最小树状设定下研究了固定温度的GRPO型估计器,基于多元德尔塔方法揭示采样维度不对称性:增加思维采样数(K)仅产生严格正的方差下界,而增加每思维的延续数(M)能使主导项方差以1/M速率趋近于零。这表明,在无价值模型的固定温度设置下,仅靠扩大思维采样无法实现精确优势估计,因此延续级分支是必要且合理的机制。实验进一步验证其有效性与必要性,在数学及视觉领域、多种模型架构与规模下均提升优化稳定性、训练效率和最终表现。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) trains Chain-of-Thought reasoning with verifiable rewards, but estimating thought-level advantages without value functions often suffers from high variance. Although tree-style branching is used in practice to reduce variance, it lacks a theoretical explanation of why it works and whether it is important or potentially necessary. We study thought-level advantage estimation in GRPO from a variance perspective under a minimal tree-style setting where multiple continuations are sampled for each thought. Using the multivariate delta method, we reveal a sampling-dimension asymmetry. Increasing sampled thoughts ($K$) leaves a strictly positive estimation-variance floor, whereas increasing continuations per thought ($M$) drives the leading-order estimation variance to zero at rate $1/M$. This implies that, within the fixed-temperature GRPO-style estimator without value models studied here, accurate thought-level advantage estimation cannot be achieved by scaling thought sampling alone, making continuation-level branching a principled and potentially necessary mechanism rather than a heuristic. Experiments further provide empirical evidence for its effectiveness and potential necessity, demonstrating improved optimization stability, training efficiency, and final performance not only in math but also across vision domains and under different model architectures and sizes.

强化学习思维链方差控制GRPO

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。