arXiv:2502.07523cs.LGcs.AI2025-02NeurIPS被引 19

提升强化学习样本效率,让训练更稳定且可扩展。

Scaling Off-Policy Reinforcement Learning with Batch and Weight Normalization

  • 在CrossQ框架中引入权重归一化,稳定高更新比例下的训练
  • 在25个复杂连续控制任务中表现优异,包括狗和人形机器人
  • 无需重置网络等复杂操作,适合追求高效稳定的RL研究者

强化学习虽已取得显著进展,但样本效率仍是实际应用的瓶颈。近期的CrossQ在极低更新/数据比(UTD=1)下实现了顶尖的样本效率。本文研究CrossQ在更高UTD比下的扩展性,发现更高更新比例会加剧训练动态问题。为此,我们向CrossQ框架引入权重归一化,该方法稳定训练过程,防止潜在可塑性丧失,并保持有效学习率恒定。所提方法在25个具有挑战性的连续控制任务上表现出色,涵盖DeepMind Control Suite和Myosuite基准,尤其在复杂环境如狗和人形机器人中表现突出。本工作无需大幅干预(如网络重置),为无模型强化学习提供了简单而稳健的样本效率与可扩展性提升路径。

原文摘要 · Abstract (English)

Reinforcement learning has achieved significant milestones, but sample efficiency remains a bottleneck for real-world applications. Recently, CrossQ has demonstrated state-of-the-art sample efficiency with a low update-to-data (UTD) ratio of 1. In this work, we explore CrossQ's scaling behavior with higher UTD ratios. We identify challenges in the training dynamics, which are emphasized by higher UTD ratios. To address these, we integrate weight normalization into the CrossQ framework, a solution that stabilizes training, has been shown to prevent potential loss of plasticity and keeps the effective learning rate constant. Our proposed approach reliably scales with increasing UTD ratios, achieving competitive performance across 25 challenging continuous control tasks on the DeepMind Control Suite and Myosuite benchmarks, notably the complex dog and humanoid environments. This work eliminates the need for drastic interventions, such as network resets, and offers a simple yet robust pathway for improving sample efficiency and scalability in model-free reinforcement learning.

强化学习样本效率稳定性

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。