arXiv:2603.12612cs.LGcs.AI2026-03被引 1

让高维人形机器人控制中的随机策略发挥更大潜力,显著提升性能。

FastDSAC: Unlocking the Potential of Maximum Entropy RL in High-Dimensional Humanoid Control

  • 分维度动态调节探索预算,改善高维环境下的探索效率。
  • 在Basketball和Balance Hard任务上分别提升180%和350%。
  • 适合追求稳定高维连续控制的科研与工程应用。

将最大熵强化学习(Maximum Entropy RL)扩展至高维人形机器人控制仍面临根本性挑战,因‘维度诅咒’导致探索效率低下和训练不稳定。因此,当前高吞吐场景主要依赖高度优化的确定性策略梯度方法。本文提出FastDSAC框架,有效释放最大熵随机策略在复杂连续控制中的潜力。引入分维度熵调节(DEM)机制,动态重分配探索预算;同时设计连续分布批评者,通过缓解高维过估计与离散量化伪影,确保价值估计准确性。在HumanoidBench及多样化连续控制任务上的大量实验表明,FastDSAC在所评估基准上建立了高维随机策略的最先进性能。该方法在多数情况下可媲美甚至超越强确定性基线,在Basketball和Balance Hard任务上分别实现180%和350%的性能提升。

原文摘要 · Abstract (English)

Scaling Maximum Entropy Reinforcement Learning (RL) to high-dimensional humanoid control remains a fundamental challenge, as the ''curse of dimensionality'' induces severe exploration inefficiency and training instability. Consequently, highly optimized deterministic policy gradients currently dominate high-throughput regimes. We address this limitation with FastDSAC, a framework that effectively unlocks the potential of maximum entropy stochastic policies for complex continuous control. We introduce Dimension-wise Entropy Modulation (DEM) to dynamically redistribute the exploration budget, alongside a continuous distributional critic tailored to ensure accurate value estimation by mitigating both high-dimensional overestimation and discrete quantization artifacts. Extensive evaluations on HumanoidBench and a diverse set of continuous control tasks demonstrate that FastDSAC establishes state-of-the-art performance for high-dimensional stochastic policies on the evaluated benchmarks. Our method is competitive with and often outperforms strong deterministic baselines, with gains of 180% and 350% on the challenging Basketball and Balance Hard tasks, respectively.

强化学习人形机器人随机策略高维控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。