arXiv:2505.13925cs.ROcs.LG2025-05NeurIPS被引 1

利用时间反转对称性提升强化学习机器人操作的样本效率

Time Reversal Symmetry for Efficient Robotic Manipulations in Deep Reinforcement Learning

  • 通过轨迹反转增强与奖励引导,实现时间对称任务的高效学习
  • 在Robosuite和MetaWorld上显著提升样本效率与最终性能
  • 适合研究高效强化学习或机器人控制的开发者参考

对称性在机器人领域普遍存在,并被广泛用于提升深度强化学习(DRL)的样本效率。然而,现有方法主要关注空间对称性(如反射、旋转、平移),而忽视了时间对称性。为填补这一空白,本文探索了时间反转对称性——一种常见于开门、关门等任务中的时序对称形式。我们提出时间反转增强的深度强化学习框架(TR-DRL),结合轨迹反转数据增强与时间反转引导的奖励塑造机制,以高效解决具有时间对称性的任务。该方法通过动态一致性过滤器识别完全可逆的转移,生成反向经验;对部分可逆转移,则根据反向任务的成功轨迹进行奖励塑形。在Robosuite与MetaWorld基准上的大量实验表明,TR-DRL在单任务与多任务设置下均有效,相比基线方法实现了更高的样本效率与更强的最终性能。

原文摘要 · Abstract (English)

Symmetry is pervasive in robotics and has been widely exploited to improve sample efficiency in deep reinforcement learning (DRL). However, existing approaches primarily focus on spatial symmetries, such as reflection, rotation, and translation, while largely neglecting temporal symmetries. To address this gap, we explore time reversal symmetry, a form of temporal symmetry commonly found in robotics tasks such as door opening and closing. We propose Time Reversal symmetry enhanced Deep Reinforcement Learning (TR-DRL), a framework that combines trajectory reversal augmentation and time reversal guided reward shaping to efficiently solve temporally symmetric tasks. Our method generates reversed transitions from fully reversible transitions, identified by a proposed dynamics-consistent filter, to augment the training data. For partially reversible transitions, we apply reward shaping to guide learning, according to successful trajectories from the reversed task. Extensive experiments on the Robosuite and MetaWorld benchmarks demonstrate that TR-DRL is effective in both single-task and multi-task settings, achieving higher sample efficiency and stronger final performance compared to baseline methods.

强化学习机器人控制对称性样本效率

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。