arXiv:2603.01302cs.RO2026-03被引 1

解决机器人操作中混合动作空间的过估计偏差问题,提升训练稳定性。

Hybrid TD3: Overestimation Bias Analysis and Stable Policy Optimization for Hybrid Action Space

  • 提出基于双批评家的混合动作空间处理方法,理论分析过估计偏差。
  • 引入加权截断Q学习目标,减少偏差并提升策略平滑性。
  • 在高维动作空间和域随机化下表现更稳定,适合复杂机器人控制任务。

在离散-连续混合动作空间中进行强化学习对机器人操作构成根本挑战,需联合优化高层任务决策与底层关节空间执行。现有方法或对连续部分进行离散化,或把离散选择松弛为连续近似,均存在高维动作空间下可扩展性差和训练不稳定的问题,尤其在域随机化条件下。本文提出 Hybrid TD3,作为 Twin Delayed Deep Deterministic Policy Gradient (TD3) 的扩展,以原则性方式原生处理参数化混合动作空间。我们对混合动作设置中的过估计偏差进行了严谨的理论分析,推导出在双批评家架构下的形式化边界,并在同步高斯误差假设下建立了五种算法变体的完整偏差排序。基于此分析,我们引入一种加权截断Q学习目标,通过对离散动作分布进行边缘化,实现与标准截断最小化相当的偏差降低,同时提升策略平滑性。实验结果表明,Hybrid TD3 在训练稳定性上优于现有先进基线方法,性能具有竞争力。

原文摘要 · Abstract (English)

Reinforcement learning in discrete-continuous hybrid action spaces presents fundamental challenges for robotic manipulation, where high-level task decisions and low-level joint-space execution must be jointly optimized. Existing approaches either discretize continuous components or relax discrete choices into continuous approximations, which suffer from scalability limitations and training instability in high-dimensional action spaces and under domain randomization. In this paper, we propose Hybrid TD3, an extension of Twin Delayed Deep Deterministic Policy Gradient (TD3) that natively handles parameterized hybrid action spaces in a principled manner. We conduct a rigorous theoretical analysis of overestimation bias in hybrid action settings, deriving formal bounds under twin-critic architectures and establishing a complete bias ordering across five algorithmic variants under synchronized Gaussian error assumptions. Building on this analysis, we introduce a weighted clipped Q-learning target that marginalizes over the discrete action distribution, achieving equivalent bias reduction to standard clipped minimization while improving policy smoothness. Experimental results demonstrate that Hybrid TD3 achieves superior training stability and competitive performance against state-of-the-art hybrid action baselines.

强化学习混合动作机器人控制策略优化

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。