arXiv:2501.13883cs.LGcs.NE2025-01被引 2

用进化策略训练基于Transformer的强化学习智能体,效果出色。

Utilizing Evolution Strategies to Train Transformers in Reinforcement Learning

  • 采用可并行的进化策略优化Transformer架构的策略网络
  • 在MuJoCo和Atari环境中均实现高性能智能体
  • 验证了进化策略对复杂模型的有效性,适合大规模模型训练

我们探索了进化策略在强化学习设置中训练基于Transformer架构策略代理的能力。通过使用OpenAI高度可并行化的进化策略,在MuJoCo Humanoid运动环境和Atari游戏环境中训练Decision Transformer,测试该黑箱优化技术对相对较大且复杂的模型(相较于文献中此前测试的模型)的训练能力。所考察的进化策略总体上表现出色,能够取得优异结果,并成功生成高性能智能体,展示了进化策略在训练此类复杂模型方面的潜力。

原文摘要 · Abstract (English)

We explore the capability of evolution strategies to train an agent with a policy based on a transformer architecture in a reinforcement learning setting. We performed experiments using OpenAI's highly parallelizable evolution strategy to train Decision Transformer in the MuJoCo Humanoid locomotion environment and in the environment of Atari games, testing the ability of this black-box optimization technique to train even such relatively large and complicated models (compared to those previously tested in the literature). The examined evolution strategy proved to be, in general, capable of achieving strong results and managed to produce high-performing agents, showcasing evolution's ability to tackle the training of even such complex models.

进化策略Transformer强化学习训练方法

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。