arXiv:2511.11607cs.LGcs.AI2025-11被引 2

通过聚类正交化提升强化学习稳定性,加快训练速度。

Clustering-Based Weight Orthogonalization for Stabilizing Deep Reinforcement Learning

  • 用聚类和投影矩阵设计新层,减少梯度干扰。
  • 在视觉和状态基任务上分别提升9%和12.6%性能。
  • 可适配多种算法,对非平稳环境有强鲁棒性。

强化学习(RL)在诸多任务中已实现超人表现,但其常假设环境平稳,而现实中环境多为非平稳,导致学习效率低下,需数百万次迭代。为此,我们提出聚类正交权重修改(COWM)层,可嵌入任意RL算法的策略网络中,有效缓解非平稳性。该层通过聚类技术和投影矩阵稳定学习过程,不仅提升训练速度,还降低梯度干扰,显著提高整体效率。实验表明,COWM在基于视觉和基于状态的DMControl基准上分别取得9%和12.6%的性能提升,且在多种算法与任务间表现出强泛化性与鲁棒性。

原文摘要 · Abstract (English)

Reinforcement learning (RL) has made significant advancements, achieving superhuman performance in various tasks. However, RL agents often operate under the assumption of environmental stationarity, which poses a great challenge to learning efficiency since many environments are inherently non-stationary. This non-stationarity results in the requirement of millions of iterations, leading to low sample efficiency. To address this issue, we introduce the Clustering Orthogonal Weight Modified (COWM) layer, which can be integrated into the policy network of any RL algorithm and mitigate non-stationarity effectively. The COWM layer stabilizes the learning process by employing clustering techniques and a projection matrix. Our approach not only improves learning speed but also reduces gradient interference, thereby enhancing the overall learning efficiency. Empirically, the COWM outperforms state-of-the-art methods and achieves improvements of 9% and 12.6% in vision based and state-based DMControl benchmark. It also shows robustness and generality across various algorithms and tasks.

强化学习稳定性聚类

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。