用不完美演示数据拼接轨迹,提升机器人模仿学习效果
Robust Offline Imitation Learning Through State-level Trajectory Stitching
- 基于状态搜索拼接低质量示范中的动作片段
- 在标准基准和真实机器人任务中显著提升性能
- 适合处理有噪声或不完整示范数据的场景
模仿学习(IL)通过专家示范使机器人掌握视觉-运动技能。但传统方法依赖高质量、稀少的专家数据,且易受协变量偏移影响。近期离线模仿学习将次优、未标注数据纳入训练以缓解此问题。本文提出一种新方法,通过利用任务相关轨迹片段和丰富的环境动态,从混合质量的离线数据中增强策略学习。具体而言,引入基于状态的搜索框架,将不完美示范中的状态-动作对拼接成更多样、信息量更高的训练轨迹。在标准模仿学习基准和真实机器人任务上的实验表明,该方法显著提升了泛化能力和性能表现。
原文摘要 · Abstract (English)
Imitation learning (IL) has proven effective for enabling robots to acquire visuomotor skills through expert demonstrations. However, traditional IL methods are limited by their reliance on high-quality, often scarce, expert data, and suffer from covariate shift. To address these challenges, recent advances in offline IL have incorporated suboptimal, unlabeled datasets into the training. In this paper, we propose a novel approach to enhance policy learning from mixed-quality offline datasets by leveraging task-relevant trajectory fragments and rich environmental dynamics. Specifically, we introduce a state-based search framework that stitches state-action pairs from imperfect demonstrations, generating more diverse and informative training trajectories. Experimental results on standard IL benchmarks and real-world robotic tasks showcase that our proposed method significantly improves both generalization and performance.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。