提出奖励感知的默认表示,让强化学习更懂奖励结构。
Reward-Aware Proto-Representations in Reinforcement Learning
- 构建奖励敏感的默认表示,从动态规划与时序差分角度学习。
- 在选项发现、探索和迁移学习中表现优于传统状态转移表示。
- 适合关注奖励结构建模的强化学习研究者使用。
近年来,后继表示(SR)在强化学习中受到越来越多关注,被用于解决探索、信用分配和泛化等关键问题。然而,标准的SR是奖励无关的,仅编码环境的转移动态。本文探讨一种新表示——默认表示(DR),它同时考虑了问题的奖励动态。我们为表格式情况下的DR建立理论基础:(1)推导动态规划方法;(2)设计时序差分学习算法;(3)刻画其向量空间基底;(4)通过默认特征将DR推广至函数逼近情形。实验分析显示,在奖励塑造、选项发现、探索与迁移学习等多个应用中,相较于SR,DR展现出定性不同的奖励感知行为,并在多个任务上取得定量性能提升。
原文摘要 · Abstract (English)
In recent years, the successor representation (SR) has attracted increasing attention in reinforcement learning (RL), and it has been used to address some of its key challenges, such as exploration, credit assignment, and generalization. The SR can be seen as representing the underlying credit assignment structure of the environment by implicitly encoding its induced transition dynamics. However, the SR is reward-agnostic. In this paper, we discuss a similar representation that also takes into account the reward dynamics of the problem. We study the default representation (DR), a recently proposed representation with limited theoretical (and empirical) analysis. Here, we lay some of the theoretical foundation underlying the DR in the tabular case by (1) deriving dynamic programming and (2) temporal-difference methods to learn the DR, (3) characterizing the basis for the vector space of the DR, and (4) formally extending the DR to the function approximation case through default features. Empirically, we analyze the benefits of the DR in many of the settings in which the SR has been applied, including (1) reward shaping, (2) option discovery, (3) exploration, and (4) transfer learning. Our results show that, compared to the SR, the DR gives rise to qualitatively different, reward-aware behaviour and quantitatively better performance in several settings.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。