从原始轨迹中自动学习奖励机,无需奖励或标签信息。
Active Reward Machine Inference From Raw State Trajectories
- 基于状态轨迹直接推断奖励机结构,无需人工标注。
- 仅需轨迹数据即可完成学习,适用于信息稀缺场景。
- 支持主动查询扩展轨迹,提升数据与计算效率。
奖励机是捕捉多阶段任务所需记忆的类自动机结构。结合强化学习或最优控制方法,可生成实现此类任务的机器人策略。然而,手动指定奖励机(包括基于高层特征决策的标记函数)是一项艰巨任务。本文研究如何直接从原始状态和策略信息中学习奖励机。不同于现有工作,我们假设无法访问奖励、标签或机器节点的观测,证明了何种轨迹数据足以在此信息匮乏条件下学习奖励机。随后,我们将结果拓展至主动学习场景,通过增量查询轨迹扩展来提升数据与计算效率。实验在多个网格世界示例中验证了方法的有效性。
原文摘要 · Abstract (English)
Reward machines are automaton-like structures that capture the memory required to accomplish a multi-stage task. When combined with reinforcement learning or optimal control methods, they can be used to synthesize robot policies to achieve such tasks. However, specifying a reward machine by hand, including a labeling function capturing high-level features that the decisions are based on, can be a daunting task. This paper deals with the problem of learning reward machines directly from raw state and policy information. As opposed to existing works, we assume no access to observations of rewards, labels, or machine nodes, and show what trajectory data is sufficient for learning the reward machine in this information-scarce regime. We then extend the result to an active learning setting where we incrementally query trajectory extensions to improve data (and indirectly computational) efficiency. Results are demonstrated with several grid world examples.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。