用可调偏置控制提升连续控制中策略学习的稳定性
Mitigating Estimation Bias with Representation Learning in TD Error-Driven Regularization
- 通过双演员-评论家框架,设计对称与非对称组合策略调节估计偏差
- 仅需一个超参数即可在过估计与欠估计间灵活切换,性能超越基准
- 适合需要精准价值估计的强化学习应用,如机器人控制
连续控制中的确定性策略梯度算法常受价值估计偏差影响而性能下降。尽管双评论家能缓解此类偏差,但双演员的探索潜力尚未被充分挖掘。本文基于时序差分误差驱动正则化(TDDR)框架,提出三种凸组合策略——对称与非对称形式,平衡悲观估计以抑制过估计,并利用双演员实现乐观探索以缓解欠估计。整个机制仅由单一超参数控制,可灵活调节偏差程度。为进一步提升性能,将增强的状态与动作表示引入演员和评论家网络。大量实验表明,该方法持续优于多个基准模型,验证了可调偏差的有效性,并揭示环境不同下过估计与欠估计可分别被差异化利用。
原文摘要 · Abstract (English)
Deterministic policy gradient algorithms for continuous control suffer from value estimation biases that degrade performance. While double critics reduce such biases, the exploration potential of double actors remains underexplored. Building on temporal-difference error-driven regularization (TDDR), a double actor-critic framework, this work introduces enhanced methods to achieve flexible bias control and stronger representation learning. We propose three convex combination strategies, symmetric and asymmetric, that balance pessimistic estimates to mitigate overestimation and optimistic exploration via double actors to alleviate underestimation. A single hyperparameter governs this mechanism, enabling tunable control across the bias spectrum. To further improve performance, we integrate augmented state and action representations into the actor and critic networks. Extensive experiments show that our approach consistently outperforms benchmarks, demonstrating the value of tunable bias and revealing that both overestimation and underestimation can be exploited differently depending on the environment.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。