Eigenoptions能同时提升探索与信用分配,但在线发现可能适得其反。
A Study of Value-Aware Eigenoptions
- 用预设的特征向量选项增强探索与信用分配
- 在线发现选项会过度影响经验,阻碍学习
- 提出非线性函数下学习选项值的方法,适合深度RL
选项通过引入时间与层次结构的归纳偏置,为强化学习提供强大框架。尽管在序列决策中表现优异,但通常为人工设计而非学习获得。在选项发现方法中,特征向量选项(eigenoptions)在探索方面表现突出,但在信用分配中的作用尚未充分研究。本文探究特征向量选项是否能加速无模型强化学习中的信用分配,在表格型与像素级网格世界中进行评估。结果表明,预设的特征向量选项不仅促进探索,还改善信用分配;而在线发现可能导致经验偏差过大,阻碍学习。在深度强化学习背景下,我们提出一种在非线性函数逼近下的选项值学习方法,强调终止条件对性能的影响。研究揭示了特征向量选项及选项框架在同时支持探索与信用分配方面的潜力与复杂性。
原文摘要 · Abstract (English)
Options, which impose an inductive bias toward temporal and hierarchical structure, offer a powerful framework for reinforcement learning (RL). While effective in sequential decision-making, they are often handcrafted rather than learned. Among approaches for discovering options, eigenoptions have shown strong performance in exploration, but their role in credit assignment remains underexplored. In this paper, we investigate whether eigenoptions can accelerate credit assignment in model-free RL, evaluating them in tabular and pixel-based gridworlds. We find that pre-specified eigenoptions aid not only exploration but also credit assignment, whereas online discovery can bias the agent's experience too strongly and hinder learning. In the context of deep RL, we also propose a method for learning option-values under non-linear function approximation, highlighting the impact of termination conditions on performance. Our findings reveal both the promise and complexity of using eigenoptions, and options more broadly, to simultaneously support credit assignment and exploration in reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。