用在线交互提升大模型模仿学习策略,稳定高效不乱动
Policy Decorator: Model-Agnostic Online Refinement for Large Policy Model
- 用残差策略在在线交互中微调大模型模仿学习
- 8个任务上显著提升性能,保持动作平滑无抖动
- 适合需要稳定高质策略的机器人应用
近期机器人学习利用大型模型和大量示范数据发展出高效策略,但受限于示范数据的数量、质量和多样性。本文探索通过与环境的在线交互来改进离线训练的模仿学习模型。提出Policy Decorator,采用模型无关的残差策略,在在线交互中对大型模仿学习模型进行优化。通过实施受控探索策略,实现稳定且样本高效的在线学习。评估涵盖两个基准(ManiSkill和Adroit)上的八个任务,涉及两种先进模仿学习模型(Behavior Transformer和Diffusion Policy)。结果表明,Policy Decorator能有效提升离线训练策略,同时保持模仿学习模型的平滑动作特性,避免纯强化学习策略的异常行为。
原文摘要 · Abstract (English)
Recent advancements in robot learning have used imitation learning with large models and extensive demonstrations to develop effective policies. However, these models are often limited by the quantity, quality, and diversity of demonstrations. This paper explores improving offline-trained imitation learning models through online interactions with the environment. We introduce Policy Decorator, which uses a model-agnostic residual policy to refine large imitation learning models during online interactions. By implementing controlled exploration strategies, Policy Decorator enables stable, sample-efficient online learning. Our evaluation spans eight tasks across two benchmarks-ManiSkill and Adroit-and involves two state-of-the-art imitation learning models (Behavior Transformer and Diffusion Policy). The results show Policy Decorator effectively improves the offline-trained policies and preserves the smooth motion of imitation learning models, avoiding the erratic behaviors of pure RL policies. See our project page (https://policydecorator.github.io) for videos.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。