arXiv:2503.04256cs.LGcs.AI2025-03被引 9

DRAGO让强化学习模型在任务切换时记住旧知识,避免遗忘。

Knowledge Retention for Continual Model-Based Reinforcement Learning

  • 用生成模型合成历史经验,无需存储原始数据。
  • 通过内在奖励引导探索旧任务相关状态,有效恢复记忆。
  • 适合需要长期累积经验的连续学习场景。

我们提出DRAGO,一种新型持续式模型基强化学习方法,旨在提升在一系列奖励函数不同但状态空间与动态不变的任务序列中世界模型的增量构建能力。DRAGO包含两个核心组件:合成经验回放,利用生成模型从过往任务生成合成经验,使智能体在不存储数据的前提下强化已学动态;通过探索重获记忆,引入内在奖励机制,引导智能体重新访问先前任务的相关状态。二者协同使智能体能够维持一个全面且持续演进的世界模型,从而在多样环境中实现更高效的学习与适应。实证评估表明,DRAGO可在多种持续学习场景中有效保留知识,显著提升性能。

原文摘要 · Abstract (English)

We propose DRAGO, a novel approach for continual model-based reinforcement learning aimed at improving the incremental development of world models across a sequence of tasks that differ in their reward functions but not the state space or dynamics. DRAGO comprises two key components: Synthetic Experience Rehearsal, which leverages generative models to create synthetic experiences from past tasks, allowing the agent to reinforce previously learned dynamics without storing data, and Regaining Memories Through Exploration, which introduces an intrinsic reward mechanism to guide the agent toward revisiting relevant states from prior tasks. Together, these components enable the agent to maintain a comprehensive and continually developing world model, facilitating more effective learning and adaptation across diverse environments. Empirical evaluations demonstrate that DRAGO is able to preserve knowledge across tasks, achieving superior performance in various continual learning scenarios.

强化学习持续学习世界模型记忆保留

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。