让强化学习采样更智能:自动保留对称性,提升机器人控制效率
Equivariant Action Sampling for Reinforcement Learning and Planning
- 基于对称性设计动作采样方法,确保决策过程符合物理规律
- 在多个连续控制任务中,性能显著优于传统随机采样方法
- 适合研究机器人控制、运动规划的学者与工程师参考
连续控制任务中的强化学习算法依赖精确的基于采样的动作选择。许多任务(如机器人操作)具有固有的对称性,但如何在采样方法中正确融入对称性仍是一大挑战。本文提出一种能保持目标对称性的动作采样方法,并应用于坐标回归问题,结果表明该对称感知采样方法远超朴素采样。进一步构建了基于对称性的模型预测路径积分(MPPI)通用框架。在多个连续控制任务上对比标准采样方法,实验验证了该方法的有效性,凸显了在采样动作选择中保持对称性的重要性。
原文摘要 · Abstract (English)
Reinforcement learning (RL) algorithms for continuous control tasks require accurate sampling-based action selection. Many tasks, such as robotic manipulation, contain inherent problem symmetries. However, correctly incorporating symmetry into sampling-based approaches remains a challenge. This work addresses the challenge of preserving symmetry in sampling-based planning and control, a key component for enhancing decision-making efficiency in RL. We introduce an action sampling approach that enforces the desired symmetry. We apply our proposed method to a coordinate regression problem and show that the symmetry aware sampling method drastically outperforms the naive sampling approach. We furthermore develop a general framework for sampling-based model-based planning with Model Predictive Path Integral (MPPI). We compare our MPPI approach with standard sampling methods on several continuous control tasks. Empirical demonstrations across multiple continuous control environments validate the effectiveness of our approach, showcasing the importance of symmetry preservation in sampling-based action selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。