通过在线模型选择提升强化学习训练效率与稳定性
Improved Training Mechanism for Reinforcement Learning via Online Model Selection
- 根据实时表现动态挑选最优强化学习代理
- 在多种任务中实现资源高效分配与训练稳定
- 适合追求训练鲁棒性的算法研发人员
我们研究强化学习中的在线模型选择问题,即选择器可访问一组强化学习智能体,并学会自适应地选择配置最合适的智能体。目标是验证将在线模型选择方法融入强化学习训练过程后,能带来的效率提升与性能增益。从理论上分析了实践中识别正确配置的有效特征,针对三个实际标准展开理论探讨:1)高效资源分配;2)非平稳动态下的自适应能力;3)不同随机种子下的训练稳定性。理论结果得到多个强化学习模型选择任务的实证支持,包括神经网络架构选择、步长选择以及自模型选择。
原文摘要 · Abstract (English)
We study the problem of online model selection in reinforcement learning, where the selector has access to a class of reinforcement learning agents and learns to adaptively select the agent with the right configuration. Our goal is to establish the improved efficiency and performance gains achieved by integrating online model selection methods into reinforcement learning training procedures. We examine the theoretical characterizations that are effective for identifying the right configuration in practice, and address three practical criteria from a theoretical perspective: 1) Efficient resource allocation, 2) Adaptation under non-stationary dynamics, and 3) Training stability across different seeds. Our theoretical results are accompanied by empirical evidence from various model selection tasks in reinforcement learning, including neural architecture selection, step-size selection, and self model selection.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。