从部分观测策略中学习奖励机器,揭示隐藏的奖励结构。
Learning Reward Machines from Partially Observed Policies
- 构建前缀树策略表征状态与命题序列的行动分布
- 在有限深度下可精确恢复奖励机器的等价类
- 适用于离散、连续及真实动物行为数据
逆强化学习旨在从最优策略或专家示范中推断奖励函数。本文假设奖励由奖励机器表示,其转移依赖于马尔可夫决策过程(MDP)状态的原子命题。目标是利用有限信息识别真实的奖励机器。为此,首先引入前缀树策略,将每个可达的原子命题有限序列与MDP状态关联,定义行动分布。接着,刻画了在给定前缀树策略下可识别的奖励机器等价类。最后提出一种基于SAT的算法,利用前缀树策略提取的信息求解奖励机器。证明若前缀树策略已知至足够(但有限)深度,该算法可恢复奖励机器的精确等价类。该充分深度为MDP状态数与奖励机器状态数上界的函数。结果进一步推广至仅能获取最优策略示范的情形。通过离散网格世界、积木世界、连续状态空间机械臂及真实小鼠实验数据验证了方法的有效性与通用性。
原文摘要 · Abstract (English)
Inverse reinforcement learning is the problem of inferring a reward function from an optimal policy or demonstrations by an expert. In this work, it is assumed that the reward is expressed as a reward machine whose transitions depend on atomic propositions associated with the state of a Markov Decision Process (MDP). Our goal is to identify the true reward machine using finite information. To this end, we first introduce the notion of a prefix tree policy which associates a distribution of actions to each state of the MDP and each attainable finite sequence of atomic propositions. Then, we characterize an equivalence class of reward machines that can be identified given the prefix tree policy. Finally, we propose a SAT-based algorithm that uses information extracted from the prefix tree policy to solve for a reward machine. It is proved that if the prefix tree policy is known up to a sufficient (but finite) depth, our algorithm recovers the exact reward machine up to the equivalence class. This sufficient depth is derived as a function of the number of MDP states and (an upper bound on) the number of states of the reward machine. These results are further extended to the case where we only have access to demonstrations from an optimal policy. Several examples, including discrete grid and block worlds, a continuous state-space robotic arm, and real data from experiments with mice, are used to demonstrate the effectiveness and generality of the approach.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。