arXiv:2508.03194cs.LG2025-08综述被引 1

系统梳理了强化学习中数据、网络和训练预算的扩展策略,为提升决策性能提供路线图。

Scaling DRL for Decision Making: A Survey on Data, Network, and Training Budget Strategies

  • 从数据、网络架构到训练资源三方面分析规模化方法
  • 揭示大规模训练与模型表现之间的正向关系
  • 适合关注强化学习高效训练的研究者和工程师

近年来,神经网络模型和训练数据的扩大推动了深度学习在计算机视觉和自然语言处理领域的显著进展。这一进步基于缩放定律:增加模型参数和训练数据可提升学习性能。尽管这些领域已取得突破(如GPT-4、Midjourney),但缩放定律在深度强化学习(DRL)中的应用仍相对有限。本文系统分析了数据、网络和训练预算三个维度的缩放策略。在数据扩展中,研究并行采样与数据生成方法,探索数据量与学习效果的关系;在网络扩展中,探讨单体扩张、集成与混合专家(MoE)方法及智能体数量扩展,提升模型表达能力但带来计算挑战;在训练预算扩展中,评估分布式训练、高重放缓存率、大批次和辅助训练对效率与收敛的影响。通过整合这些策略,本文不仅揭示其协同作用,还为未来研究提供方向,强调在可扩展性与计算效率间平衡的重要性,并指出利用缩放潜力在机器人控制、自动驾驶和大模型训练等任务中的前景。

原文摘要 · Abstract (English)

In recent years, the expansion of neural network models and training data has driven remarkable progress in deep learning, particularly in computer vision and natural language processing. This advancement is underpinned by the concept of Scaling Laws, which demonstrates that scaling model parameters and training data enhances learning performance. While these fields have witnessed breakthroughs, such as the development of large language models like GPT-4 and advanced vision models like Midjourney, the application of scaling laws in deep reinforcement learning (DRL) remains relatively unexplored. Despite its potential to improve performance, the integration of scaling laws into DRL for decision making has not been fully realized. This review addresses this gap by systematically analyzing scaling strategies in three dimensions: data, network, and training budget. In data scaling, we explore methods to optimize data efficiency through parallel sampling and data generation, examining the relationship between data volume and learning outcomes. For network scaling, we investigate architectural enhancements, including monolithic expansions, ensemble and MoE methods, and agent number scaling techniques, which collectively enhance model expressivity while posing unique computational challenges. Lastly, in training budget scaling, we evaluate the impact of distributed training, high replay ratios, large batch sizes, and auxiliary training on training efficiency and convergence. By synthesizing these strategies, this review not only highlights their synergistic roles in advancing DRL for decision making but also provides a roadmap for future research. We emphasize the importance of balancing scalability with computational efficiency and outline promising directions for leveraging scaling to unlock the full potential of DRL in various tasks such as robot control, autonomous driving and LLM training.

强化学习规模扩展决策制定

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。