arXiv:2503.01837cs.LGcs.CV2025-03ICML被引 10

用少量示范提升机器人长序列操作的训练效率

Multi-Stage Manipulation with Demonstration-Augmented Reward, Policy, and World Model Learning

  • 结合多阶段奖励与世界模型,从视觉输入中学习分步目标
  • 仅需5次示范即在16个任务上实现平均40%的数据效率提升
  • 适合低数据场景下的复杂机器人控制,尤其擅长视觉任务

机器人长序列操作任务在强化学习中面临巨大挑战,主要源于难以设计密集奖励函数以及在广阔状态-动作空间中有效探索。然而,这些任务通常具有多阶段结构,可被用于将整体目标分解为可管理的子目标。本文提出DEMO3框架,利用这一结构实现从视觉输入高效学习。方法融合多阶段密集奖励学习、双阶段训练策略和世界模型学习,构建了一个示范增强的强化学习框架,显著缓解了长时序任务中的探索难题。评估表明,相比现有最优方法,该方法在平均上提升数据效率40%,在特别困难的任务上提升达70%。实验覆盖16个稀疏奖励任务,涵盖四个领域,包括仅用5次示范即可完成的人形机器人视觉控制任务。

原文摘要 · Abstract (English)

Long-horizon tasks in robotic manipulation present significant challenges in reinforcement learning (RL) due to the difficulty of designing dense reward functions and effectively exploring the expansive state-action space. However, despite a lack of dense rewards, these tasks often have a multi-stage structure, which can be leveraged to decompose the overall objective into manageable subgoals. In this work, we propose DEMO3, a framework that exploits this structure for efficient learning from visual inputs. Specifically, our approach incorporates multi-stage dense reward learning, a bi-phasic training scheme, and world model learning into a carefully designed demonstration-augmented RL framework that strongly mitigates the challenge of exploration in long-horizon tasks. Our evaluations demonstrate that our method improves data-efficiency by an average of 40% and by 70% on particularly difficult tasks compared to state-of-the-art approaches. We validate this across 16 sparse-reward tasks spanning four domains, including challenging humanoid visual control tasks using as few as five demonstrations.

机器人操作强化学习少样本学习世界模型

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。