提出新框架与适配方法,让机器人在复杂任务中成功组合行为策略。
Composing Option Sequences by Adaptation: Initial Results
- 通过三种适配策略调整策略起点与终点,提升序列成功率。
- 实验表明,五步深度强化学习策略组合在未适配时几乎无法完成任务。
- 适合研究机器人自主决策与策略组合的学者参考。
现实世界中的机器人操作常需根据当前情境调整行为,例如改变策略执行顺序以完成目标。然而,我们发现即使启动与终止条件匹配,由五个深度强化学习选项组成的新型序列也难以成功完成拾取-放置任务。为此,本文提出一个先验判断序列是否成功的框架,并研究三种使不成功序列得以有效执行的适配方法:(1) 训练第二项策略从第一项结束位置开始;(2) 训练第一项策略到达第二项起始位置的中心点;(3) 训练第一项策略到达第二项起始位置的中位点。实验结果表明,该框架及适配方法在使策略适应新序列方面具有显著潜力。
原文摘要 · Abstract (English)
Robot manipulation in real-world settings often requires adapting the robot's behavior to the current situation, such as by changing the sequences in which policies execute to achieve the desired task. Problematically, however, we show that composing a novel sequence of five deep RL options to perform a pick-and-place task is unlikely to successfully complete, even if their initiation and termination conditions align. We propose a framework to determine whether sequences will succeed a priori, and examine three approaches that adapt options to sequence successfully if they will not. Crucially, our adaptation methods consider the actual subset of points that the option is trained from or where it ends: (1) trains the second option to start where the first ends; (2) trains the first option to reach the centroid of where the second starts; and (3) trains the first option to reach the median of where the second starts. Our results show that our framework and adaptation methods have promise in adapting options to work in novel sequences.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。