通过多组对称性增强,提升机器人操作的采样效率。
Multi-Group Equivariant Augmentation for Reinforcement Learning in Robot Manipulation
- 引入非等距对称性,在时空维度上独立应用多组变换。
- 在仿真与真实机器人上验证,显著提升离线强化学习采样效率。
- 适合关注机器人视觉运动学习与数据效率的研究者。
视觉运动学习在真实机器人操作中的部署高度依赖采样效率。尽管任务对称性作为有效的归纳偏置已被证明能提升效率,但现有方法大多局限于等距对称性——即在所有时间步对所有任务物体施加相同的群变换。本文探索非等距对称性,允许在空间和时间维度上独立应用多组变换以放宽限制。我们提出了一种包含非等距对称结构的新形式部分可观察马尔可夫决策过程(POMDP)模型,并设计了一种简单有效的数据增强方法:多组等变增强(MEA)。将MEA与离线强化学习结合,并引入基于体素的视觉表示以保持平移等变性。在两个操作领域中的大量仿真与真实机器人实验表明该方法有效。
原文摘要 · Abstract (English)
Sampling efficiency is critical for deploying visuomotor learning in real-world robotic manipulation. While task symmetry has emerged as a promising inductive bias to improve efficiency, most prior work is limited to isometric symmetries -- applying the same group transformation to all task objects across all timesteps. In this work, we explore non-isometric symmetries, applying multiple independent group transformations across spatial and temporal dimensions to relax these constraints. We introduce a novel formulation of the partially observable Markov decision process (POMDP) that incorporates the non-isometric symmetry structures, and propose a simple yet effective data augmentation method, Multi-Group Equivariance Augmentation (MEA). We integrate MEA with offline reinforcement learning to enhance sampling efficiency, and introduce a voxel-based visual representation that preserves translational equivariance. Extensive simulation and real-robot experiments across two manipulation domains demonstrate the effectiveness of our approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。