arXiv:2410.00064cs.LGcs.AI2024-10ICRA被引 14

M2Distill通过多模态知识蒸馏,解决机械臂持续学习中的技能遗忘问题。

M2Distill: Multi-Modal Distillation for Lifelong Imitation Learning

  • 采用多模态蒸馏,保持视觉、语言、动作分布的潜在空间一致
  • 在LIBERO多个基准上优于现有方法,任务成功率显著提升
  • 适合需要长期积累新技能且不丢旧技能的应用场景

针对机械操作任务中持续学习导致的分布偏移问题,现有方法常依赖无监督技能发现或多重策略蒸馏,存在可扩展性差及潜在空间不一致的问题,易引发已学技能的灾难性遗忘。本文提出M2Distill,一种基于多模态蒸馏的持续模仿学习方法,通过调控不同模态(视觉、语言、动作)在前后学习步骤间的潜在表示变化,并减少连续步骤间高斯混合模型(GMM)策略的差异,确保所学策略在保留旧任务能力的同时无缝融合新技能。在LIBERO终身模仿学习基准套件(包括LIBERO-OBJECT、LIBERO-GOAL和LIBERO-SPATIAL)上的大量实验表明,该方法在所有评估指标上均持续优于现有最先进方法。

原文摘要 · Abstract (English)

Lifelong imitation learning for manipulation tasks poses significant challenges due to distribution shifts that occur in incremental learning steps. Existing methods often focus on unsupervised skill discovery to construct an ever-growing skill library or distillation from multiple policies, which can lead to scalability issues as diverse manipulation tasks are continually introduced and may fail to ensure a consistent latent space throughout the learning process, leading to catastrophic forgetting of previously learned skills. In this paper, we introduce M2Distill, a multi-modal distillation-based method for lifelong imitation learning focusing on preserving consistent latent space across vision, language, and action distributions throughout the learning process. By regulating the shifts in latent representations across different modalities from previous to current steps, and reducing discrepancies in Gaussian Mixture Model (GMM) policies between consecutive learning steps, we ensure that the learned policy retains its ability to perform previously learned tasks while seamlessly integrating new skills. Extensive evaluations on the LIBERO lifelong imitation learning benchmark suites, including LIBERO-OBJECT, LIBERO-GOAL, and LIBERO-SPATIAL, demonstrate that our method consistently outperforms prior state-of-the-art methods across all evaluated metrics.

持续学习模仿学习多模态蒸馏机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。