arXiv:2411.04466cs.LGcs.AI2024-11NeurIPS被引 4

用进化方法自动生成多样化任务,让智能体在复杂模拟环境中高效自适应训练。

Enabling Adaptive Agent Training in Open-Ended Simulators by Targeting Diversity

  • 提出进化式任务生成框架DIVA,无需预设参数即可自动设计多样任务
  • 在复杂开放环境上训练的智能体表现远超现有基线方法
  • 结合无监督与有监督优势,适合真实场景模拟中的自适应学习

端到端学习在具身决策领域应用受限于对大量代表性训练数据的依赖。元强化学习(meta-RL)放弃零样本泛化目标,转向少样本适应,有望缩小泛化差距。尽管学习这种元级自适应行为仍需大量数据,但日益复杂的高效仿真器正被广泛应用。然而,为这些复杂领域手动设计足够多样且数量充足的训练任务代价高昂。域随机化(DR)和过程生成(PG)虽可缓解此问题,但要求仿真器具备能直接映射任务多样性的精细参数——这一假设同样不切实际。本文提出DIVA,一种针对复杂、开放仿真环境的进化式任务生成方法。与无监督环境设计(UED)类似,DIVA可应用于任意参数化,同时还能融入现实可用的领域知识,兼具UED的灵活性与DR/PG对良好设计仿真器结构的利用。实验表明,DIVA能有效应对复杂参数化,成功训练出自适应智能体,显著优于现有文献中的竞争基线。结果凸显了此类半监督环境设计(SSED)方法的潜力,而DIVA是该方向首个奠基性工作,有望推动真实仿真域中的智能体训练,提升其鲁棒性与能力。

原文摘要 · Abstract (English)

The wider application of end-to-end learning methods to embodied decision-making domains remains bottlenecked by their reliance on a superabundance of training data representative of the target domain. Meta-reinforcement learning (meta-RL) approaches abandon the aim of zero-shot generalization--the goal of standard reinforcement learning (RL)--in favor of few-shot adaptation, and thus hold promise for bridging larger generalization gaps. While learning this meta-level adaptive behavior still requires substantial data, efficient environment simulators approaching real-world complexity are growing in prevalence. Even so, hand-designing sufficiently diverse and numerous simulated training tasks for these complex domains is prohibitively labor-intensive. Domain randomization (DR) and procedural generation (PG), offered as solutions to this problem, require simulators to possess carefully-defined parameters which directly translate to meaningful task diversity--a similarly prohibitive assumption. In this work, we present DIVA, an evolutionary approach for generating diverse training tasks in such complex, open-ended simulators. Like unsupervised environment design (UED) methods, DIVA can be applied to arbitrary parameterizations, but can additionally incorporate realistically-available domain knowledge--thus inheriting the flexibility and generality of UED, and the supervised structure embedded in well-designed simulators exploited by DR and PG. Our empirical results showcase DIVA's unique ability to overcome complex parameterizations and successfully train adaptive agent behavior, far outperforming competitive baselines from prior literature. These findings highlight the potential of such semi-supervised environment design (SSED) approaches, of which DIVA is the first humble constituent, to enable training in realistic simulated domains, and produce more robust and capable adaptive agents.

强化学习任务生成仿真训练自适应智能体

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。