arXiv:2412.10096cs.ROcs.LG2024-12被引 4

从视觉示范中自动学习机器人操作的奖励机器结构。

Reward Machine Inference for Robotic Manipulation

  • 直接从视觉演示中联合学习奖励机器结构与关键事件
  • 无需预设命题或稀疏奖励信号,准确捕捉任务时序结构
  • 适合无先验知识的复杂操作任务学习,提升强化学习效率

基于演示的学习(LfD)和强化学习(RL)使机器人能够完成复杂任务。奖励机器(RMs)通过组织高层任务信息,增强强化学习在长时程任务中的能力。本文提出一种新型的LfD方法,可直接从机器人操作的视觉示范中学习奖励机器。与以往方法不同,该方法无需预定义命题或对稀疏奖励信号的先验知识,而是联合学习奖励机器结构,并识别驱动状态转移的关键高阶事件。我们在基于视觉的操纵任务上验证了该方法,结果表明所推断的奖励机器能准确捕获任务结构,并使强化学习代理有效学习到最优策略。

原文摘要 · Abstract (English)

Learning from Demonstrations (LfD) and Reinforcement Learning (RL) have enabled robot agents to accomplish complex tasks. Reward Machines (RMs) enhance RL's capability to train policies over extended time horizons by structuring high-level task information. In this work, we introduce a novel LfD approach for learning RMs directly from visual demonstrations of robotic manipulation tasks. Unlike previous methods, our approach requires no predefined propositions or prior knowledge of the underlying sparse reward signals. Instead, it jointly learns the RM structure and identifies key high-level events that drive transitions between RM states. We validate our method on vision-based manipulation tasks, showing that the inferred RM accurately captures task structure and enables an RL agent to effectively learn an optimal policy.

机器人操作奖励机器模仿学习强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。