arXiv:2506.05702cs.LGcs.AI2025-06被引 1

提出动态动作空间下的持续学习框架,实现策略泛化

Action-Adaptive Continual Learning: Enabling Policy Generalization under Dynamic Action Spaces

  • 用动作表征空间解耦策略与动作空间,支持动态适应
  • 在新动作空间下自适应微调编码器-解码器,平衡稳定与可塑性
  • 构建新基准验证,适合需要跨场景泛化的强化学习应用

持续学习(CL)使智能体能够按顺序学习任务,积累过往知识并用于未来问题求解或学习。然而,现有方法通常假设智能体能力在动态环境中保持不变,这与现实情况不符。本文提出一个更真实的问题:动态能力下的持续学习(CL-DC),核心挑战在于如何在不同动作空间间实现策略泛化。受大脑皮层功能启发,我们提出动作自适应持续学习框架(AACL)。该框架通过构建动作表征空间,将智能体策略与具体动作空间解耦。面对新动作空间时,动作表征的编码器-解码器被自适应微调,以维持稳定性与可塑性的平衡。此外,我们基于三个环境构建了一个新基准,用于验证CL-DC方法的有效性。实验结果表明,本框架在跨动作空间策略泛化方面显著优于主流方法。

原文摘要 · Abstract (English)

Continual Learning (CL) is a powerful tool that enables agents to learn a sequence of tasks, accumulating knowledge learned in the past and using it for problem-solving or future task learning. However, existing CL methods often assume that the agent's capabilities remain static within dynamic environments, which doesn't reflect real-world scenarios where capabilities dynamically change. This paper introduces a new and realistic problem: Continual Learning with Dynamic Capabilities (CL-DC), posing a significant challenge for CL agents: How can policy generalization across different action spaces be achieved? Inspired by the cortical functions, we propose an Action-Adaptive Continual Learning framework (AACL) to address this challenge. Our framework decouples the agent's policy from the specific action space by building an action representation space. For a new action space, the encoder-decoder of action representations is adaptively fine-tuned to maintain a balance between stability and plasticity. Furthermore, we release a benchmark based on three environments to validate the effectiveness of methods for CL-DC. Experimental results demonstrate that our framework outperforms popular methods by generalizing the policy across action spaces.

持续学习策略泛化动态动作空间强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。