用低成本代理模型加速强化学习训练,提升动态系统决策效率。
Accelerating Reinforcement Learning Training Using Simulation Surrogate Models

- 构建仿真代理模型近似复杂系统输入输出关系
- 实验显示训练速度提升显著,支持动态环境重训练
- 适合需要快速迭代的实时决策场景
高保真仿真模型广泛用于分析复杂随机系统,但其高昂计算成本推动了低成本代理模型的发展,以近似原仿真模型的输入-输出关系。与此同时,强化学习(RL)已成为在随机环境中进行在线决策的强大框架,越来越多研究将仿真模型作为训练环境。本文研究适用于奖励结构、模型参数或系统动态随时间变化场景的代理模型,探索其与仿真模型及强化学习模型的交互机制。通过基于离散事件仿真的随机服务系统数值实验,结果表明利用代理模型可显著加速强化学习的训练与再训练过程。
原文摘要 · Abstract (English)
High-fidelity simulation models are widely used to analyze complex stochastic systems, but their high computational cost motivates the development of cheaper surrogate models that approximate the simulation model's input-output relationship. In parallel, reinforcement learning (RL) has emerged as a powerful framework for making online decisions in stochastic environments, with increasing attention being given to the use of simulation models as training environments for RL models. We investigate a class of surrogate models suitable for accelerating RL training in settings where the reward structure, model parameters, or system dynamics change over time and explore their interactions with simulation models and RL models. Through numerical experiments on a stochastic service system modeled via discrete-event simulation, we demonstrate that leveraging surrogate models can substantially accelerate RL training and re-training.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。