arXiv:2511.12706cs.LGcs.AI2025-11AAAI被引 1

让任务和关卡自动匹配,训练更高效的强化学习智能体

Beyond Fixed Tasks: Unsupervised Environment Design for Task-Level Pairs

  • 联合生成任务与关卡的自动课程,确保可解且具挑战性
  • 在难以采样可解组合时,性能显著优于随机采样方法
  • 利用任务和关卡结构的突变机制,加速策略收敛

在复杂环境中训练通用智能体执行复杂指令(任务)仍是强化学习的核心挑战。随机采样任务-关卡对常导致无法求解的组合,凸显了任务与关卡协同设计的必要性。尽管无监督环境设计(UED)已证明能自动构建关卡课程,但以往工作仅针对固定任务。本文提出ATLAS(Aligning Tasks and Levels for Autocurricula of Specifications),一种在任务与关卡上联合生成自动课程的新方法。该方法基于UED,自动产生可解且具挑战性的任务-关卡对以供策略训练。为评估ATLAS并推动领域进展,我们引入一个评测套件,将任务建模为奖励机,在Minigrid环境中进行测试。实验表明,当可解组合难以采样时,ATLAS显著优于随机采样方法。此外,利用任务与关卡结构的突变机制,能加速收敛至高性能策略。

原文摘要 · Abstract (English)

Training general agents to follow complex instructions (tasks) in intricate environments (levels) remains a core challenge in reinforcement learning. Random sampling of task-level pairs often produces unsolvable combinations, highlighting the need to co-design tasks and levels. While unsupervised environment design (UED) has proven effective at automatically designing level curricula, prior work has only considered a fixed task. We present ATLAS (Aligning Tasks and Levels for Autocurricula of Specifications), a novel method that generates joint autocurricula over tasks and levels. Our approach builds upon UED to automatically produce solvable yet challenging task-level pairs for policy training. To evaluate ATLAS and drive progress in the field, we introduce an evaluation suite that models tasks as reward machines in Minigrid levels. Experiments demonstrate that ATLAS vastly outperforms random sampling approaches, particularly when sampling solvable pairs is unlikely. We further show that mutations leveraging the structure of both tasks and levels accelerate convergence to performant policies.

强化学习自动课程任务设计关卡生成

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。