将GRPO扩展到连续控制,为机器人强化学习提供理论框架
Extending Group Relative Policy Optimization to Continuous Control: A Theoretical Framework for Robotic Reinforcement Learning
- 用轨迹聚类和状态感知优势估计解决连续动作难题
- 理论证明收敛性与计算复杂度,适用于行走与抓取任务
- 适合研究机器人强化学习的学者和工程师
Group Relative Policy Optimization (GRPO) 在离散动作空间中表现出色,通过基于分组的优势估计消除了对价值函数的依赖。然而,其在连续控制中的应用尚未探索,限制了其在需要连续动作的机器人领域的应用。本文提出一个理论框架,将GRPO扩展至连续控制环境,解决了高维动作空间、稀疏奖励和时序动态等挑战。方法包括基于轨迹的策略聚类、状态感知的优势估计以及针对机器人应用设计的正则化策略更新。我们提供了收敛性与计算复杂度的理论分析,为未来在运动与操作任务中的实证验证奠定基础。
原文摘要 · Abstract (English)
Group Relative Policy Optimization (GRPO) has shown promise in discrete action spaces by eliminating value function dependencies through group-based advantage estimation. However, its application to continuous control remains unexplored, limiting its utility in robotics where continuous actions are essential. This paper presents a theoretical framework extending GRPO to continuous control environments, addressing challenges in high-dimensional action spaces, sparse rewards, and temporal dynamics. Our approach introduces trajectory-based policy clustering, state-aware advantage estimation, and regularized policy updates designed for robotic applications. We provide theoretical analysis of convergence properties and computational complexity, establishing a foundation for future empirical validation in robotic systems including locomotion and manipulation tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。