利用对称性提升强化学习的样本效率与泛化能力
Equivariant Goal Conditioned Contrastive Reinforcement Learning
- 引入旋转不变的评判器和旋转变换的智能体,构建对称约束
- 在多种模拟任务中,性能优于主流基线方法
- 适用于需要空间泛化的机器人操作场景
对比强化学习(CRL)为从无标签交互中提取有用结构化表征提供了有前景的框架。通过拉近状态-动作对与其对应未来状态的距离,同时推开负样本对,CRL 可在无需人工设计奖励的情况下学习非平凡策略。本文提出等变对比强化学习(ECRL),进一步利用目标条件化任务中的内在对称性来结构化潜在空间。我们形式化定义了目标条件化群不变马尔可夫决策过程(Goal-Conditioned Group-Invariant MDPs),用于刻画具有旋转对称性的机器人操作任务,并在此基础上引入一种新型旋转不变的评判器与旋转等变的智能体,应用于对比强化学习。该方法在基于状态和图像的多种模拟任务中持续优于强基线。最后,我们将方法扩展至离线强化学习设置,在多个任务上验证了其有效性。
原文摘要 · Abstract (English)
Contrastive Reinforcement Learning (CRL) provides a promising framework for extracting useful structured representations from unlabeled interactions. By pulling together state-action pairs and their corresponding future states, while pushing apart negative pairs, CRL enables learning nontrivial policies without manually designed rewards. In this work, we propose Equivariant CRL (ECRL), which further structures the latent space using equivariant constraints. By leveraging inherent symmetries in goal-conditioned manipulation tasks, our method improves both sample efficiency and spatial generalization. Specifically, we formally define Goal-Conditioned Group-Invariant MDPs to characterize rotation-symmetric robotic manipulation tasks, and build on this by introducing a novel rotation-invariant critic representation paired with a rotation-equivariant actor for Contrastive RL. Our approach consistently outperforms strong baselines across a range of simulated tasks in both state-based and image-based settings. Finally, we extend our method to the offline RL setting, demonstrating its effectiveness across multiple tasks.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。