arXiv:2512.10365cs.LGcs.AI2025-12
提出适用于Transformer策略的广义策略梯度理论,统一现有方法。
GPG: Generalized Policy Gradient Theorem for Transformer-based Policies
- 基于Transformer架构构建广义策略梯度框架
- 证明标准策略梯度与GRPO均为其特例
- 为大模型高效策略优化提供新思路,适合强化学习研究者
我们提出了专为基于Transformer的策略设计的广义策略梯度(GPG)定理。值得注意的是,我们证明了标准策略梯度定理和GRPO均是该GPG框架下的特例。此外,我们探索了该理论在训练大语言模型(LLMs)中的实际应用,为高效的策略优化提供了新的见解。
原文摘要 · Abstract (English)
We present the Generalized Policy Gradient (GPG) Theorem, specifically designed for Transformer-based policies. Notably, we demonstrate that both standard Policy Gradient Theorem and GRPO emerge as special cases within our GPG framework. Furthermore, we explore its practical applications in training Large Language Models (LLMs), offering new insights into efficient policy optimization.
强化学习Transformer策略梯度大模型
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。