用准蒙特卡洛初始化提升元强化学习训练速度
Quasi-Monte Carlo Initialization for Meta-Reinforcement Learning

- 采用准蒙特卡洛方法初始化元学习权重,优化搜索效率
- 在相似环境中收敛速度比标准正交初始化快
- 适合需要快速适应新连续控制任务的研究者
本文探讨了在现代基准环境中的元强化学习中,使用准蒙特卡洛(QMC)权重初始化的有效性。通过多种采样方法对基于种群的搜索进行约束,并从一组基础任务中聚合出最优先验。当外推到相似的未见连续控制环境时,QMC元先验在训练收敛性上优于现代正交(SB3)默认初始化。而在任务差异较大的情况下,正交初始化在无偏搜索中表现更优。
原文摘要 · Abstract (English)
This paper explores the efficacy of quasi-Monte Carlo (QMC) weight initialization for meta-reinforcement learning within modern benchmark environments. Various sampling methods are used to bound a population-based search and aggregate an optimal prior from a baseline set of tasks. The QMC meta-priors show improvements in training convergence compared to modern orthogonal (SB3) defaults when extrapolated to similar unseen continuous control environments. In dissimilar tasks, the orthogonal orientation was globally superior for an unbiased search.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。