arXiv:2510.14065cs.RO2025-10被引 4

将强化学习技能融入任务与运动规划,提升复杂场景下的执行效率。

Optimistic Reinforcement Learning-Based Skill Insertions for Task and Motion Planning

  • 用数据驱动逻辑组件定义可被符号规划调用的强化学习技能
  • 引入计划优化子流程,缓解不确定性带来的执行偏差
  • 在概率性动作场景中显著提升规划效率,适合机器人长时序操作

机器人操作中的任务与运动规划(TAMP)需要长时序推理,涉及多种动作与技能。尽管确定性动作可通过采样或约束优化生成,但带有不确定性的概率动作在TAMP中仍具挑战。相反,强化学习(RL)擅长获取多样且鲁棒的短时序操作技能。本文提出一种将RL技能集成到TAMP流程中的方法:除策略外,还通过数据驱动的逻辑组件定义技能,使其可被符号规划调用;同时设计计划优化子流程以应对不确定性影响。实验对比了来自TAMP和RL领域的基线层级规划方法,结果表明,嵌入RL技能后,TAMP可扩展至包含概率技能的领域,并在规划效率上优于以往方法。

原文摘要 · Abstract (English)

Task and motion planning (TAMP) for robotics manipulation necessitates long-horizon reasoning involving versatile actions and skills. While deterministic actions can be crafted by sampling or optimizing with certain constraints, planning actions with uncertainty, i.e., probabilistic actions, remains a challenge for TAMP. On the contrary, Reinforcement Learning (RL) excels in acquiring versatile, yet short-horizon, manipulation skills that are robust with uncertainties. In this letter, we design a method that integrates RL skills into TAMP pipelines. Besides the policy, a RL skill is defined with data-driven logical components that enable the skill to be deployed by symbolic planning. A plan refinement sub-routine is designed to further tackle the inevitable effect uncertainties. In the experiments, we compare our method with baseline hierarchical planning from both TAMP and RL fields and illustrate the strength of the method. The results show that by embedding RL skills, we extend the capability of TAMP to domains with probabilistic skills, and improve the planning efficiency compared to the previous methods.

强化学习任务规划机器人操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。