arXiv:2510.09369cs.CL2025-10被引 4

通过马尔可夫似然链接组级奖励与词元,提升大模型推理稳定性。

Token-Level Policy Optimization: Linking Group-Level Rewards to Token-Level Aggregation via Markov Likelihood

  • 用序列似然将组级奖励映射到词元级别,实现精准优化
  • 在数学推理任务上超越基线,准确率与@k指标全面领先
  • 解决熵崩溃问题,适合需要稳定训练的复杂推理场景

组相对策略优化(GRPO)显著提升了大语言模型(LLMs)的推理能力,尤其在数学任务上表现突出。然而,GRPO及其相关熵正则化方法仍受链式思维(CoT)中稀疏词元奖励的制约。现有方法常采用统一的词元级熵调整,易导致熵崩溃或模型崩溃。本文提出TEPO,一种新颖的词元级框架,通过马尔可夫似然(序列似然)将组级奖励与词元通过词元级聚合相连接。实验表明,TEPO在关键指标(包括@k和准确率)上持续优于现有基线,不仅在数学推理任务上达到新SOTA,还显著提升了训练稳定性。

原文摘要 · Abstract (English)

Group Relative Policy Optimization (GRPO) has significantly advanced the reasoning ability of large language models (LLMs), particularly by boosting their mathematical performance. However, GRPO and related entropy-regularization methods still face challenges rooted in the sparse token rewards inherent to chain-of-thought (CoT). Current approaches often rely on undifferentiated token-level entropy adjustments, which frequently lead to entropy collapse or model collapse. In this work, we propose TEPO, a novel token-level framework that incorporates Markov Likelihood (sequence likelihood) links group-level rewards with tokens via token-level aggregation. Experiments show that TEPO consistently outperforms existing baselines across key metrics (including @k and accuracy). It not only sets a new state of the art on mathematical reasoning tasks but also significantly enhances training stability.

大模型推理策略优化强化学习数学推理

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。