通过马尔可夫似然链接组级奖励与词元,提升大模型推理稳定性。
Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Markov Likelihood
- 用序列似然将组级奖励映射到词元级别,实现精准优化
- 在数学推理任务上超越基线,准确率与@k指标全面领先
- 解决熵崩溃问题,适合需要稳定训练的复杂推理场景
组相对策略优化(GRPO)显著提升了大语言模型(LLMs)的推理能力,尤其在数学任务上表现突出。然而,GRPO及其相关熵正则化方法仍受链式思维(CoT)中稀疏词元奖励的制约。现有方法常采用统一的词元级熵调整,易导致熵崩溃或模型崩溃。本文提出TEPO,一种新颖的词元级框架,通过马尔可夫似然(序列似然)将组级奖励与词元通过词元级聚合相连接。实验表明,TEPO在关键指标(包括@k和准确率)上持续优于现有基线,不仅在数学推理任务上达到新SOTA,还显著提升了训练稳定性。
原文摘要 · Abstract (English)
Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly by boosting their mathematical performance. However, GRPO and related entropy-regularization methods still face challenges rooted in the sparse token rewards inherent to chain-of-thought (CoT). Current approaches often rely on undifferentiated token-level entropy adjustments, which frequently lead to entropy collapse or model collapse. In this work, we propose TEPO, a novel token-level framework that incorporates Markov Likelihood (sequence likelihood) links group-level rewards with tokens via token-level aggregation. Experiments show that TEPO consistently outperforms existing baselines across key metrics (including @k and accuracy). It not only sets a new state of the art on mathematical reasoning tasks but also significantly enhances training stability.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。