用预测编码改进决策变压器,提升复杂任务下的决策能力
Predictive Coding for Decision Transformer
- 引入预测编码作为未来条件,增强决策的时序建模能力
- 在8个基准数据集上表现优于或媲美主流方法
- 适合长时序、稀疏奖励的离线目标导向任务
近期离线强化学习研究证明,将决策问题建模为回报条件化的监督学习是有效的。特别是决策变压器(DT)架构在多个领域展现出潜力。然而,尽管初期成功,DT在若干具有挑战性的目标条件化强化学习数据集上表现不佳。这一局限源于回报条件化在引导策略学习上的低效性,尤其在非结构化和次优数据集上,导致DT难以有效学习时序组合性。该问题在长时程稀疏奖励任务中可能进一步加剧。为此,我们提出预测编码决策变压器(PCDT)框架,利用广义未来条件化增强DT方法。PCDT扩展了DT架构,以预测编码为条件,使决策同时基于过去和未来因素,从而提升泛化性能。在AntMaze和FrankaKitchen环境的8个数据集上,本方法表现达到或超过现有主流基于价值和基于Transformer的方法。此外,我们在一个物理机器人目标达成任务上也进行了评估。
原文摘要 · Abstract (English)
Recent work in offline reinforcement learning (RL) has demonstrated the effectiveness of formulating decision-making as return-conditioned supervised learning. Notably, the decision transformer (DT) architecture has shown promise across various domains. However, despite its initial success, DTs have underperformed on several challenging datasets in goal-conditioned RL. This limitation stems from the inefficiency of return conditioning for guiding policy learning, particularly in unstructured and suboptimal datasets, resulting in DTs failing to effectively learn temporal compositionality. Moreover, this problem might be further exacerbated in long-horizon sparse-reward tasks. To address this challenge, we propose the Predictive Coding for Decision Transformer (PCDT) framework, which leverages generalized future conditioning to enhance DT methods. PCDT utilizes an architecture that extends the DT framework, conditioned on predictive codings, enabling decision-making based on both past and future factors, thereby improving generalization. Through extensive experiments on eight datasets from the AntMaze and FrankaKitchen environments, our proposed method achieves performance on par with or surpassing existing popular value-based and transformer-based methods in offline goal-conditioned RL. Furthermore, we also evaluate our method on a goal-reaching task with a physical robot.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。