arXiv:2503.03145cs.RO2025-03被引 2

用因果关系指导强化学习,让机器人更高效完成多阶段任务

Causality-Based Reinforcement Learning Method for Multi-Stage Robotic Tasks

  • 通过发现动作与奖励的因果关系,构建仅包含有效动作的空间
  • 减少探索冗余和进程倒退,提升多阶段任务成功率
  • 适合需要分步执行的复杂机器人任务研究者

深度强化学习在各类机器人任务中取得了显著进展,但在处理多阶段任务时仍面临挑战。强化学习算法常出现冗余探索、陷入死胡同和进程倒退等问题。为此,我们提出一种将因果关系融入强化学习的方法。该方法使机器人能够自动发现动作与任务奖励之间的因果关系,并仅基于因果动作构建动作空间,从而减少冗余探索和进程倒退。通过使用因果策略梯度法将正确的因果关系整合进学习过程,该方法可显著提升强化学习算法在多阶段机器人任务中的性能。

原文摘要 · Abstract (English)

Deep reinforcement learning has made significant strides in various robotic tasks. However, employing deep reinforcement learning methods to tackle multi-stage tasks still a challenge. Reinforcement learning algorithms often encounter issues such as redundant exploration, getting stuck in dead ends, and progress reversal in multi-stage tasks. To address this, we propose a method that integrates causal relationships with reinforcement learning for multi-stage tasks. Our approach enables robots to automatically discover the causal relationships between their actions and the rewards of the tasks and constructs the action space using only causal actions, thereby reducing redundant exploration and progress reversal. By integrating correct causal relationships using the causal policy gradient method into the learning process, our approach can enhance the performance of reinforcement learning algorithms in multi-stage robotic tasks.

强化学习因果推理机器人任务

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。