arXiv:2502.02869cs.LGcs.AI2025-02NeurIPS被引 4

用随机生成任务提升大规模上下文强化学习的泛化能力

Towards Large-Scale In-Context Reinforcement Learning by Meta-Training in Randomized Worlds

  • 设计随机生成的马尔可夫决策过程AnyMDP,批量构建高质量训练任务
  • 在大规模任务集上训练后,模型能适应未见任务并实现跨场景泛化
  • 研究揭示多样性与适应时长的权衡,指导更鲁棒的ICRL系统设计

上下文强化学习(ICRL)使智能体能够从交互经验中自动、实时地学习。然而,其规模化面临任务集合难以扩展的挑战。为此,我们提出程序化生成的表格型马尔可夫决策过程——AnyMDP。通过精心设计的随机化机制,AnyMDP可在大规模下生成高质量任务,同时保持较低的结构偏差。为支持高效的大规模元训练,我们引入解耦策略蒸馏,并在ICRL框架中注入先验信息。结果表明,在足够大规模的AnyMDP任务集上训练后,所提模型可通过多样化的上下文学习范式泛化到训练集中未包含的任务。AnyMDP提供的可扩展任务集还支持对数据分布与ICRL性能关系的更深入实证分析。我们进一步发现,ICRL的泛化能力可能以增加任务多样性与更长适应周期为代价。这一发现对构建稳健的规模化ICRL系统具有重要意义,强调了多样化、大范围任务设计的必要性,并应优先考虑长期性能而非少样本适应。

原文摘要 · Abstract (English)

In-Context Reinforcement Learning (ICRL) enables agents to learn automatically and on-the-fly from their interactive experiences. However, a major challenge in scaling up ICRL is the lack of scalable task collections. To address this, we propose the procedurally generated tabular Markov Decision Processes, named AnyMDP. Through a carefully designed randomization process, AnyMDP is capable of generating high-quality tasks on a large scale while maintaining relatively low structural biases. To facilitate efficient meta-training at scale, we further introduce decoupled policy distillation and induce prior information in the ICRL framework. Our results demonstrate that, with a sufficiently large scale of AnyMDP tasks, the proposed model can generalize to tasks that were not considered in the training set through versatile in-context learning paradigms. The scalable task set provided by AnyMDP also enables a more thorough empirical investigation of the relationship between data distribution and ICRL performance. We further show that the generalization of ICRL potentially comes at the cost of increased task diversity and longer adaptation periods. This finding carries critical implications for scaling robust ICRL capabilities, highlighting the necessity of diverse and extensive task design, and prioritizing asymptotic performance over few-shot adaptation.

强化学习上下文学习任务生成元学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。