arXiv:2409.19231cs.LGcs.AI2024-09被引 8

提出双演员-评论家框架,用时序差分误差正则化提升价值估计精度。

Double Actor-Critic with TD Error-Driven Regularization in Reinforcement Learning

  • 采用双演员-评论家结构,每演员配独立评论家。
  • 引入时序差分误差驱动的正则化机制,提升估计稳定性。
  • 无需新增超参数,适合复杂连续控制任务应用。

为提升强化学习中的价值估计性能,我们提出一种基于双演员-评论家框架并结合时序差分误差驱动正则化的新型算法,简称TDDR。TDDR采用双演员结构,每个演员与一个评论家配对,充分借鉴双评论家的优势。此外,TDDR设计了一种创新的评论家正则化架构。相比传统基于确定性策略梯度的算法(缺乏双演员-评论家结构),TDDR在价值估计上表现更优。同时,相较于现有双演员-评论家框架算法,TDDR未引入任何额外超参数,显著简化了算法设计与实现过程。实验表明,在具有挑战性的连续控制任务中,TDDR相较于基准算法展现出强大竞争力。

原文摘要 · Abstract (English)

To obtain better value estimation in reinforcement learning, we propose a novel algorithm based on the double actor-critic framework with temporal difference error-driven regularization, abbreviated as TDDR. TDDR employs double actors, with each actor paired with a critic, thereby fully leveraging the advantages of double critics. Additionally, TDDR introduces an innovative critic regularization architecture. Compared to classical deterministic policy gradient-based algorithms that lack a double actor-critic structure, TDDR provides superior estimation. Moreover, unlike existing algorithms with double actor-critic frameworks, TDDR does not introduce any additional hyperparameters, significantly simplifying the design and implementation process. Experiments demonstrate that TDDR exhibits strong competitiveness compared to benchmark algorithms in challenging continuous control tasks.

强化学习双演员价值估计

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。