arXiv:2502.02834cs.LGcs.AI2025-02ICML被引 2

通过虚拟训练提升元强化学习在分布外任务上的泛化能力

Task-Aware Virtual Training: Enhancing Generalization in Meta-Reinforcement Learning for Out-of-Distribution Tasks

  • 用度量学习构建任务表征,捕捉真实与虚拟任务特征
  • 在多种MuJoCo和MetaWorld环境中显著提升分布外任务表现
  • 适合研究元学习泛化与鲁棒性优化的学者参考

元强化学习旨在训练出能泛化到未见任务的策略,这些任务来自任务分布。尽管基于上下文的元强化学习方法通过任务潜变量改进任务表征,但通常在分布外(OOD)任务上表现不佳。为此,我们提出任务感知虚拟训练(TAVT),一种新算法,利用基于度量的学习方法准确捕捉训练及分布外场景下的任务特征。该方法在虚拟任务中有效保持任务特性,并采用状态正则化技术减轻状态变化环境中的过估计误差。数值实验表明,TAVT在多个MuJoCo和MetaWorld环境中显著增强了对分布外任务的泛化能力。代码已公开于https://github.com/JM-Kim-94/tavt.git。

原文摘要 · Abstract (English)

Meta reinforcement learning aims to develop policies that generalize to unseen tasks sampled from a task distribution. While context-based meta-RL methods improve task representation using task latents, they often struggle with out-of-distribution (OOD) tasks. To address this, we propose Task-Aware Virtual Training (TAVT), a novel algorithm that accurately captures task characteristics for both training and OOD scenarios using metric-based representation learning. Our method successfully preserves task characteristics in virtual tasks and employs a state regularization technique to mitigate overestimation errors in state-varying environments. Numerical results demonstrate that TAVT significantly enhances generalization to OOD tasks across various MuJoCo and MetaWorld environments. Our code is available at https://github.com/JM-Kim-94/tavt.git.

元强化学习泛化能力分布外虚拟训练

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。