通过引入力矩动作空间,提升四足机器人强化学习的探索能力。
WARL: Wrench-Augmented Reinforcement Learning for Task-Agnostic Learning in Legged Robots

- 在动作空间中加入力矩(wrench),增强早期探索能力。
- 可在多种地形和任务上稳定训练,无需调整奖励或复杂课程设计。
- 适合希望提升强化学习效率的机器人控制研究者。
尽管强化学习在四足机器人运动控制中已取得优异性能,但其受限于关节空间动作的探索能力不足。为此,本文提出一种新方法——力矩增强强化学习(WARL),将力和力矩(wrench)引入动作空间。该方法结合力矩引导探索与基于成功率的课程机制,在学习初期显著扩展探索范围,最终实现仅依赖关节控制的行为获取。实验使用四足机器人验证,WARL能在多种地形和运动任务下稳健学习,无需针对地形调整奖励函数或设计复杂课程。消融实验表明,渐进式移除力矩的切换课程有效提升性能。同时发现,过度引入力矩可能导致行为未能充分利用机器人物理特性。结果表明,力矩增强探索能有效提升学习效率,但需与机器人本体结构相匹配,这是未来关键挑战。
原文摘要 · Abstract (English)
While reinforcement learning for legged robots has achieved high motor performance, it has been constrained by the limited exploration capability of actions confined to the joint space. To address this issue, this study proposes a new method, Wrench-Augmented Reinforcement Learning (WARL), which introduces a wrenche (force and torque) into the action space. The proposed method combines wrench-guided exploration with a success rate-based curriculum mechanism to expand exploration capabilities in the early stages of learning, with the ultimate goal of acquiring behaviors based solely on joint control. Experiments using a quadruped robot demonstrated that WARL can learn robustly across diverse terrains and motor tasks without requiring terrain-specific reward adjustments or complex curriculum designs. Furthermore, an ablation study verified the effectiveness of the Switching Curriculum, which gradually eliminates the wrench. On the other hand, we also show that introducing a wrench can encourage behaviors that do not sufficiently exploit the robot's physical embodiment. These findings suggest that while wrench-based exploration enhancement is effective for improving learning efficiency, designing it in a way that is consistent with the robot's physical structure is a critical future challenge.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。