arXiv:2411.00361cs.LG2024-11被引 2

用偏好优化解决分层强化学习中的目标不可达与策略不稳问题

Direct Preference Optimization for Primitive-Enabled Hierarchical RL: A Bilevel Approach

  • 将分层强化学习建模为双层优化,通过偏好比较训练高层策略
  • 在机器人导航与操作任务中提升40%性能,显著减少无效目标生成
  • 适合研究分层决策、目标规划及偏好学习的从业者参考

分层强化学习(HRL)通过将复杂长周期任务分解为可管理的子任务,使智能体能够解决高难度问题。然而,现有方法面临两大挑战:(i) 低层策略在训练过程中动态变化导致高层学习不稳定;(ii) 高层策略生成无法实现的子目标。为此,我们提出DIPPER,一种新的分层强化学习框架,将带目标条件的HRL建模为双层优化问题,并利用直接偏好优化(DPO)训练高层策略。通过学习关于子目标序列的静态偏好比较,而非依赖不断变化的奖励信号,DIPPER有效缓解了非平稳性对分层学习的影响。为应对不可达子目标问题,DIPPER引入低层价值函数正则化,促使高层策略生成可实现的目标。我们还设计了两项新指标,定量验证DIPPER在缓解非平稳性和子目标不可行性方面的有效性。在具有挑战性的机器人导航和操作基准测试中,DIPPER相比现有最优基线性能提升最高达40%,证明基于偏好的方法能有效解决分层强化学习中的长期难题。

原文摘要 · Abstract (English)

Hierarchical reinforcement learning (HRL) enables agents to solve complex, long-horizon tasks by decomposing them into manageable sub-tasks. However, HRL methods face two fundamental challenges: (i) non-stationarity caused by the evolving lower-level policy during training, which destabilizes higher-level learning, and (ii) the generation of infeasible subgoals that lower-level policies cannot achieve. To address these challenges, we introduce DIPPER, a novel HRL framework that formulates goal-conditioned HRL as a bi-level optimization problem and leverages direct preference optimization (DPO) to train the higher-level policy. By learning from stationary preference comparisons over subgoal sequences rather than rewards that depend on the evolving lower-level policy, DIPPER mitigates the impact of non-stationarity on hierarchical learning. To address infeasible subgoals, DIPPER incorporates lower-level value function regularization that encourages the higher-level policy to propose achievable subgoals. We also introduce two novel metrics to quantitatively verify that DIPPER mitigates non-stationarity and infeasible subgoal generation issues in HRL. We perform empirical evaluations on challenging robotic navigation and manipulation benchmarks and show that DIPPER achieves upto 40% improvements over state-of-the-art baselines, demonstrating that preference-based methods can effectively alleviate persistent challenges in hierarchical

分层RL偏好优化机器人控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。