arXiv:2503.14809cs.LGcs.AI2025-03

用专家抽象指导学习,让机器人更高效地完成多个连续控制任务。

Learning with Expert Abstractions for Efficient Multi-Task Continuous Control

  • 基于专家提供的高层次任务抽象动态规划生成子目标
  • 稀疏奖励下通过抽象模型最优状态价值重塑奖励,提升学习效率
  • 支持零样本泛化,适合复杂多任务连续控制场景

在复杂连续多任务环境中,决策常因难以获取精确的规划模型,以及纯试错学习效率低下而受阻。尽管环境动态难以精确建模,但人类专家往往能提供高保真的抽象,捕捉任务的核心结构与用户偏好。现有层次化方法多针对离散场景,且难以跨任务泛化。本文提出一种层次强化学习方法:通过在专家指定的抽象空间中动态规划,生成子目标来训练目标条件策略。为应对稀疏奖励挑战,我们基于抽象模型中的最优状态价值进行奖励塑造。该结构化决策流程提升了样本效率,并促进零样本泛化。在一系列程序生成的连续控制环境上的实证评估表明,该方法在样本效率、任务完成率、复杂任务扩展性及新场景泛化能力上均优于现有层次强化学习方法。

原文摘要 · Abstract (English)

Decision-making in complex, continuous multi-task environments is often hindered by the difficulty of obtaining accurate models for planning and the inefficiency of learning purely from trial and error. While precise environment dynamics may be hard to specify, human experts can often provide high-fidelity abstractions that capture the essential high-level structure of a task and user preferences in the target environment. Existing hierarchical approaches often target discrete settings and do not generalize across tasks. We propose a hierarchical reinforcement learning approach that addresses these limitations by dynamically planning over the expert-specified abstraction to generate subgoals to learn a goal-conditioned policy. To overcome the challenges of learning under sparse rewards, we shape the reward based on the optimal state value in the abstract model. This structured decision-making process enhances sample efficiency and facilitates zero-shot generalization. Our empirical evaluation on a suite of procedurally generated continuous control environments demonstrates that our approach outperforms existing hierarchical reinforcement learning methods in terms of sample efficiency, task completion rate, scalability to complex tasks, and generalization to novel scenarios.

强化学习多任务连续控制抽象建模

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。