提出优化框架提升强化学习环境设计的鲁棒性。
An Optimisation Framework for Unsupervised Environment Design
- 从优化视角构建非凸强凹目标函数,理论更扎实。
- 在多种难度环境中表现优于已有方法,提升泛化能力。
- 适合关注强化学习鲁棒性与自动环境生成的研究者。
为使强化学习智能体在高风险场景中部署,必须具备对陌生情况的强鲁棒性。无监督环境设计(UED)是一类旨在最大化智能体在不同环境配置下泛化能力的方法。本文从优化角度研究UED,相比以往工作在实际场景中提供了更强的理论保障。以往方法仅在收敛时保证效果,而本框架采用非凸-强凹目标函数,在零和设定下给出了可证明收敛的算法。实验验证了该方法的有效性,在多个不同难度的环境中均优于现有方法。
原文摘要 · Abstract (English)
For reinforcement learning agents to be deployed in high-risk settings, they must achieve a high level of robustness to unfamiliar scenarios. One method for improving robustness is unsupervised environment design (UED), a suite of methods aiming to maximise an agent's generalisability across configurations of an environment. In this work, we study UED from an optimisation perspective, providing stronger theoretical guarantees for practical settings than prior work. Whereas previous methods relied on guarantees if they reach convergence, our framework employs a nonconvex-strongly-concave objective for which we provide a provably convergent algorithm in the zero-sum setting. We empirically verify the efficacy of our method, outperforming prior methods in a number of environments with varying difficulties.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。