arXiv:2601.14957cs.LG2026-01NeurIPS被引 1

提出新方法提升无监督环境生成的挑战度识别能力

Improving Regret Approximation for Unsupervised Dynamic Environment Generation

  • 设计动态环境生成机制,增强奖励信号密度
  • 提出MNA新指标,更准确识别高难度训练关卡
  • 适合研究强化学习训练集自动生成的学者

无监督环境设计(UED)旨在为强化学习(RL)代理自动构建训练课程,以提升泛化能力和零样本性能。然而,在某些环境下,仅少数参数组合就会显著增加策略复杂度,导致有效课程设计困难。现有方法面临信用分配难题,且依赖的遗憾近似无法准确识别挑战性水平,且随着环境规模增大问题加剧。本文提出动态环境生成方法(DEGen),通过增强奖励信号密度降低信用分配难度,使UED可扩展至更大规模环境。同时引入新的遗憾近似指标——最大负优势(MNA),能更精准识别更具挑战性的训练水平。实验证明,MNA优于现有近似方法,结合DEGen后在各类环境中均表现更优,尤其在环境规模增大时优势显著。代码已公开:https://github.com/HarryMJMead/Dynamic-Environment-Generation-for-UED。

原文摘要 · Abstract (English)

Unsupervised Environment Design (UED) seeks to automatically generate training curricula for reinforcement learning (RL) agents, with the goal of improving generalisation and zero-shot performance. However, designing effective curricula remains a difficult problem, particularly in settings where small subsets of environment parameterisations result in significant increases in the complexity of the required policy. Current methods struggle with a difficult credit assignment problem and rely on regret approximations that fail to identify challenging levels, both of which are compounded as the size of the environment grows. We propose Dynamic Environment Generation for UED (DEGen) to enable a denser level generator reward signal, reducing the difficulty of credit assignment and allowing for UED to scale to larger environment sizes. We also introduce a new regret approximation, Maximised Negative Advantage (MNA), as a significantly improved metric to optimise for, that better identifies more challenging levels. We show empirically that MNA outperforms current regret approximations and when combined with DEGen, consistently outperforms existing methods, especially as the size of the environment grows. We have made all our code available here: https://github.com/HarryMJMead/Dynamic-Environment-Generation-for-UED.

强化学习环境生成无监督学习课程学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。