arXiv:2502.06919cs.LGcs.AI2025-02ICLR

提出分维动作重复机制,提升连续控制的灵活性与效率

Select before Act: Spatially Decoupled Action Repetition for Continuous Control

  • 对每个动作维度独立判断是否重复,实现空间解耦
  • 在多个任务上样本效率提升,动作波动降低30%以上
  • 适合需要精细动作控制的机器人场景

强化学习在机器人操作和运动等连续控制任务中取得显著进展。不同于传统每步决策的方法,近期研究引入动作重复机制,提升了动作持续性,改善了采样效率并获得更优性能。然而,现有方法将所有动作维度整体处理,忽视各维度差异,导致决策僵化,影响策略灵活性和有效性。本文提出一种新型重复框架SDAR,通过为每个动作维度单独执行闭环“动作-重复”选择,实现空间解耦的动作重复。该设计使重复策略更具灵活性,平衡了动作持续性与多样性。在多种连续控制场景下的实验表明,相比现有框架,SDAR具备更高的采样效率、更优的策略性能以及更低的动作波动(平均减少30%以上)。结果验证了所提空间解耦重复设计的有效性。

原文摘要 · Abstract (English)

Reinforcement Learning (RL) has achieved remarkable success in various continuous control tasks, such as robot manipulation and locomotion. Different to mainstream RL which makes decisions at individual steps, recent studies have incorporated action repetition into RL, achieving enhanced action persistence with improved sample efficiency and superior performance. However, existing methods treat all action dimensions as a whole during repetition, ignoring variations among them. This constraint leads to inflexibility in decisions, which reduces policy agility with inferior effectiveness. In this work, we propose a novel repetition framework called SDAR, which implements Spatially Decoupled Action Repetition through performing closed-loop act-or-repeat selection for each action dimension individually. SDAR achieves more flexible repetition strategies, leading to an improved balance between action persistence and diversity. Compared to existing repetition frameworks, SDAR is more sample efficient with higher policy performance and reduced action fluctuation. Experiments are conducted on various continuous control scenarios, demonstrating the effectiveness of spatially decoupled repetition design proposed in this work.

强化学习连续控制动作重复机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。