arXiv:2411.05614cs.LGcs.AI2024-11中稿 · ACM Computing Surv…综述被引 27

综述深度强化学习并行分布式加速方法,助力高效训练。

Acceleration for Deep Reinforcement Learning using Parallel and Distributed Computing: A Survey

  • 构建并行分布式训练的分类体系,涵盖仿真与计算并行
  • 对比16个开源平台,评估其对快速开发的支持能力
  • 梳理前沿方向与未解难题,为研究者指明未来路径

近年来,深度强化学习在人工智能领域取得重大突破。随着训练所需的经验数据量和神经网络规模持续增长,利用并行与分布式计算加速训练过程、降低时间开销已成为迫切需求。本文系统调研基于并行与分布式计算的深度强化学习训练加速方法,全面总结当前主流技术与核心文献。提出包含学习系统架构、仿真并行、计算并行、分布式同步机制及深度进化强化学习在内的分类体系,并讨论新兴趋势与开放问题。进一步对比16个现有开源库与平台,评估其对快速开发的支持程度。最后,展望值得深入探索的未来方向。

原文摘要 · Abstract (English)

Deep reinforcement learning has led to dramatic breakthroughs in the field of artificial intelligence for the past few years. As the amount of rollout experience data and the size of neural networks for deep reinforcement learning have grown continuously, handling the training process and reducing the time consumption using parallel and distributed computing is becoming an urgent and essential desire. In this paper, we perform a broad and thorough investigation on training acceleration methodologies for deep reinforcement learning based on parallel and distributed computing, providing a comprehensive survey in this field with state-of-the-art methods and pointers to core references. In particular, a taxonomy of literature is provided, along with a discussion of emerging topics and open issues. This incorporates learning system architectures, simulation parallelism, computing parallelism, distributed synchronization mechanisms, and deep evolutionary reinforcement learning. Further, we compare 16 current open-source libraries and platforms with criteria of facilitating rapid development. Finally, we extrapolate future directions that deserve further research.

强化学习并行计算分布式训练综述

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。