用强化学习训练的智能体轨迹提升复杂仿真环境的代理模型采样效率。
Building surrogate models using trajectories of agents trained by Reinforcement Learning
- 利用强化学习训练的智能体探索高熵状态区域,生成高效采样数据。
- 混合随机、专家与熵最大化智能体的数据集在所有测试中表现最优。
- 适合需要高效采样复杂仿真系统的强化学习优化任务。
在计算成本高昂的仿真环境中,样本效率是代理建模的常见挑战。现有方法在状态空间庞大的仿真环境中效果有限。为此,我们提出一种新方法:利用强化学习训练的策略高效采样确定性仿真环境。通过与拉丁超立方采样、主动学习及克里金插值对比,对多种采样策略进行交叉验证。结果表明,包含随机智能体、专家智能体以及针对状态转移分布最大熵区域探索的智能体所生成的混合数据集,在所有数据集上均表现最佳,这对实现有意义的状态空间表征至关重要。研究证明该方法优于当前最先进水平,为在复杂模拟器上应用代理辅助的强化学习策略优化铺平了道路。
原文摘要 · Abstract (English)
Sample efficiency in the face of computationally expensive simulations is a common concern in surrogate modeling. Current strategies to minimize the number of samples needed are not as effective in simulated environments with wide state spaces. As a response to this challenge, we propose a novel method to efficiently sample simulated deterministic environments by using policies trained by Reinforcement Learning. We provide an extensive analysis of these surrogate-building strategies with respect to Latin-Hypercube sampling or Active Learning and Kriging, cross-validating performances with all sampled datasets. The analysis shows that a mixed dataset that includes samples acquired by random agents, expert agents, and agents trained to explore the regions of maximum entropy of the state transition distribution provides the best scores through all datasets, which is crucial for a meaningful state space representation. We conclude that the proposed method improves the state-of-the-art and clears the path to enable the application of surrogate-aided Reinforcement Learning policy optimization strategies on complex simulators.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。