arXiv:2512.10235cs.RO2025-12

用情境奖励机分解抓取任务,提升机器人学习效率与成功率。

Task-Oriented Grasping Using Reinforcement Learning with a Contextual Reward Machine

  • 将抓取任务拆解为带情境的子任务,动态调整奖励与动作空间。
  • 模拟中成功率达95%,真实机器人任务成功率达83.3%。
  • 适合需要高效学习和复杂任务分解的机器人抓取场景。

本文提出一种结合情境奖励机的强化学习框架,用于任务导向抓取。情境奖励机通过将抓取任务分解为具有特定上下文的子任务,降低任务复杂度。每个子任务包含阶段专属的奖励函数、动作空间和状态抽象函数,实现高效阶段内引导,并缩小状态-动作空间,明确探索边界。引入阶段间转移奖励,激励或惩罚阶段转换,引导模型学习理想的任务序列,加速收敛。该方法与近端策略优化算法结合,在1000次模拟抓取任务中取得95%的成功率,覆盖多种物体、可操作性及抓取拓扑结构,优于现有最优方法。在真实机器人上,60次抓取任务中成功率达83.3%,涉及六种可操作性。实验结果表明该模型具备更高精度、数据效率和学习效率,适用于仿真与现实世界中的任务导向抓取。

原文摘要 · Abstract (English)

This paper presents a reinforcement learning framework that incorporates a Contextual Reward Machine for task-oriented grasping. The Contextual Reward Machine reduces task complexity by decomposing grasping tasks into manageable sub-tasks. Each sub-task is associated with a stage-specific context, including a reward function, an action space, and a state abstraction function. This contextual information enables efficient intra-stage guidance and improves learning efficiency by reducing the state-action space and guiding exploration within clearly defined boundaries. In addition, transition rewards are introduced to encourage or penalize transitions between stages which guides the model toward desirable stage sequences and further accelerates convergence. When integrated with the Proximal Policy Optimization algorithm, the proposed method achieved a 95% success rate across 1,000 simulated grasping tasks encompassing diverse objects, affordances, and grasp topologies. It outperformed the state-of-the-art methods in both learning speed and success rate. The approach was transferred to a real robot, where it achieved a success rate of 83.3% in 60 grasping tasks over six affordances. These experimental results demonstrate superior accuracy, data efficiency, and learning efficiency. They underscore the model's potential to advance task-oriented grasping in both simulated and real-world settings.

机器人抓取强化学习任务分解真实部署

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。