用非专家数据提升机器人模仿学习的鲁棒性
Using Non-Expert Data to Robustify Imitation Learning via Offline Reinforcement Learning
- 用离线强化学习改造模仿学习,让其能利用非专家数据
- 在真实场景下,成功范围扩大近两倍,初始条件容忍度显著提升
- 适合做机器人实操训练,尤其数据难获取时
模仿学习在机器人复杂任务训练中表现优异,但依赖高质量、任务特定的数据,难以适应真实世界多样的物体配置和场景。相比之下,非专家数据(如游戏数据、低质量示范、部分完成任务或次优策略的采样)覆盖更广且收集成本低,但传统模仿学习无法有效利用。本文提出,通过合理设计,离线强化学习可作为工具,挖掘非专家数据价值。实验表明,在稀疏数据条件下,标准离线强化学习无效,但经简单算法改进后,无需额外假设即可有效利用这些数据。通过扩展策略分布支持范围,融合离线强化学习的模仿学习在操作任务中展现出更强的恢复能力和泛化性能,成功率在更广泛的初始条件下保持稳定。该方法可整合所有采集数据,包括不完整或次优示范,显著提升任务导向策略表现。结果凸显了算法设计对机器人鲁棒策略学习的重要性。
原文摘要 · Abstract (English)
Imitation learning has proven effective for training robots to perform complex tasks from expert human demonstrations. However, it remains limited by its reliance on high-quality, task-specific data, restricting adaptability to the diverse range of real-world object configurations and scenarios. In contrast, non-expert data -- such as play data, suboptimal demonstrations, partial task completions, or rollouts from suboptimal policies -- can offer broader coverage and lower collection costs. However, conventional imitation learning approaches fail to utilize this data effectively. To address these challenges, we posit that with right design decisions, offline reinforcement learning can be used as a tool to harness non-expert data to enhance the performance of imitation learning policies. We show that while standard offline RL approaches can be ineffective at actually leveraging non-expert data under the sparse data coverage settings typically encountered in the real world, simple algorithmic modifications can allow for the utilization of this data, without significant additional assumptions. Our approach shows that broadening the support of the policy distribution can allow imitation algorithms augmented by offline RL to solve tasks robustly, showing considerably enhanced recovery and generalization behavior. In manipulation tasks, these innovations significantly increase the range of initial conditions where learned policies are successful when non-expert data is incorporated. Moreover, we show that these methods are able to leverage all collected data, including partial or suboptimal demonstrations, to bolster task-directed policy performance. This underscores the importance of algorithmic techniques for using non-expert data for robust policy learning in robotics. Website: https://uwrobotlearning.github.io/RISE-offline/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。