从理论证明结果监督不比过程监督更难,可能改变算法设计思路。
Do We Need to Verify Step by Step? Rethinking Process Supervision from a Theoretical Perspective
- 提出轨迹测度变换引理,连接结果与步骤级监督
- 在标准假设下,结果监督的统计难度仅多出多项式因子
- 验证器可用作最优过程奖励模型,实现两种监督互通
随着大语言模型的发展,区分过程监督与结果监督——复杂推理任务中两种关键的强化学习方法——变得至关重要。尽管过程监督在长期信用分配上具有直观优势,但二者间的精确关系仍不明晰。传统观点认为结果监督因轨迹覆盖问题而更困难,导致大量投入收集细粒度过程数据。本文通过理论分析推进该争议解决:主定理表明,在标准数据覆盖假设下,结果监督的统计难度与过程监督相当,仅多出关于时域的多项式因子。核心在于提出新的轨迹测度变换引理,将基于回报的轨迹测度与步骤级分布偏移相联系。此外,在具备验证器或回放能力的场景中,证明任意策略的优势函数可作为最优过程奖励模型,建立结果与过程监督的直接关联。这些发现表明,若存在性能差距,其根源更可能是算法局限而非固有统计难度,可能重塑强化学习中的数据收集与算法设计范式。
原文摘要 · Abstract (English)
As large language models have evolved, it has become crucial to distinguish between process supervision and outcome supervision -- two key reinforcement learning approaches to complex reasoning tasks. While process supervision offers intuitive advantages for long-term credit assignment, the precise relationship between these paradigms has remained an open question. Conventional wisdom suggests that outcome supervision is fundamentally more challenging due to the trajectory-level coverage problem, leading to significant investment in collecting fine-grained process supervision data. In this paper, we take steps towards resolving this debate. Our main theorem shows that, under standard data coverage assumptions, reinforcement learning through outcome supervision is no more statistically difficult than through process supervision, up to polynomial factors in horizon. At the core of this result lies the novel Change of Trajectory Measure Lemma -- a technical tool that bridges return-based trajectory measure and step-level distribution shift. Furthermore, for settings with access to a verifier or a rollout capability, we prove that any policy's advantage function can serve as an optimal process reward model, providing a direct connection between outcome and process supervision. These findings suggest that the empirically observed performance gap -- if any -- between outcome and process supervision likely stems from algorithmic limitations rather than inherent statistical difficulties, potentially transforming how we approach data collection and algorithm design for reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。