用梦想家模型生成轨迹,提升在线决策变压器的决策效率。
DODT: Enhanced Online Decision Transformer Learning through Dreamer's Actor-Critic Trajectory Forecasting
- 用梦想家算法生成前瞻轨迹,增强决策变压器的上下文理解。
- 在多个基准上实现更高样本效率和奖励最大化。
- 适合研究模型基于强化学习与轨迹预测融合的学者。
强化学习的进步催生了能够处理复杂决策任务的先进模型。然而,高效整合世界模型与决策变压器仍是挑战。本文提出一种新方法,结合梦想家算法生成前瞻轨迹的能力与在线决策变压器的自适应学习优势。该方法支持并行训练,由梦想家生成的轨迹增强变压器的上下文决策能力,形成双向增强回路。我们在一系列具有挑战性的基准上实证验证了该方法的有效性,在样本效率和奖励最大化方面显著优于现有方法。结果表明,所提出的集成框架不仅加速学习过程,还在多样且动态的场景中展现出强鲁棒性,标志着模型基于强化学习的重要进展。
原文摘要 · Abstract (English)
Advancements in reinforcement learning have led to the development of sophisticated models capable of learning complex decision-making tasks. However, efficiently integrating world models with decision transformers remains a challenge. In this paper, we introduce a novel approach that combines the Dreamer algorithm's ability to generate anticipatory trajectories with the adaptive learning strengths of the Online Decision Transformer. Our methodology enables parallel training where Dreamer-produced trajectories enhance the contextual decision-making of the transformer, creating a bidirectional enhancement loop. We empirically demonstrate the efficacy of our approach on a suite of challenging benchmarks, achieving notable improvements in sample efficiency and reward maximization over existing methods. Our results indicate that the proposed integrated framework not only accelerates learning but also showcases robustness in diverse and dynamic scenarios, marking a significant step forward in model-based reinforcement learning.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。