arXiv:2412.01114cs.LG2024-12被引 2

用先验数据和少量示范,自动合成密集奖励,加速复杂控制任务学习。

Dense Dynamics-Aware Reward Synthesis: Integrating Prior Experience with Demonstrations

  • 融合通用先验数据与少量专家示范,生成密集动态感知奖励
  • 在长时序任务中使学习速度提升数倍,有效引导智能体抵达远距离目标
  • 适合需要快速适应新环境的连续控制场景

许多连续控制问题可建模为稀疏奖励强化学习任务。理论上,在线强化学习方法能自主探索状态空间以解决新任务,但随着任务时长增加,找到能获得非零奖励的动作序列变得指数级困难。手动设计奖励可加速特定任务的学习,但过程繁琐且需为每个新环境重复。本文提出一种系统性奖励塑造框架,从两个来源提取信息:1)与任务无关的先验数据集;2)少量特定任务的专家示范,并利用这些先验知识为当前任务合成密集的、动态感知的奖励。实验表明,该监督显著加速学习过程,分析进一步证明该方法能有效引导在线学习智能体到达遥远目标。

原文摘要 · Abstract (English)

Many continuous control problems can be formulated as sparse-reward reinforcement learning (RL) tasks. In principle, online RL methods can automatically explore the state space to solve each new task. However, discovering sequences of actions that lead to a non-zero reward becomes exponentially more difficult as the task horizon increases. Manually shaping rewards can accelerate learning for a fixed task, but it is an arduous process that must be repeated for each new environment. We introduce a systematic reward-shaping framework that distills the information contained in 1) a task-agnostic prior data set and 2) a small number of task-specific expert demonstrations, and then uses these priors to synthesize dense dynamics-aware rewards for the given task. This supervision substantially accelerates learning in our experiments, and we provide analysis demonstrating how the approach can effectively guide online learning agents to faraway goals.

强化学习奖励塑造连续控制

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。