arXiv:2603.08476cs.RO2026-03被引 3

用隐空间对齐路由让机器人从无标注演示中学会分技能操作

LAR-MoE: Latent-Aligned Routing for Mixture of Experts in Robotic Imitation Learning

  • 先学观测与未来动作的联合隐表示,再用它指导专家路由
  • 在LIBERO上用1.5亿参数达到95.2%成功率,无需标注阶段
  • 零样本迁移到猪组织手术任务,适合无监督技能分解场景

模仿学习使机器人能从示范中获取操作技能,但面对动态差异大的任务时,模型常会平均不同行为模式。混合专家(MoE)架构通过激活特定子网络来解决此问题,但需有意义的技能分解以实现专家路由。本文提出隐空间对齐路由的混合专家(LAR-MoE),采用两阶段框架,将无监督技能发现与策略学习解耦。预训练阶段通过师生协同训练,学习观测与未来动作的联合隐表示;后训练阶段,通过约束专家路由遵循学习到的隐空间结构,防止专家坍缩并保持参数效率。我们在仿真和真实硬件上评估该方法。在LIBERO基准上,使用1.5亿参数达到95.2%的平均成功率。在腹腔镜肠管抓取与牵拉任务中,无需任何阶段标注即可匹配有监督的MoE基线,并实现零样本迁移至离体猪组织。结果表明,隐空间对齐路由为无监督技能分解提供了原则性替代方案,支持从无标签示范中实现结构化专家专精。

原文摘要 · Abstract (English)

Imitation learning enables robots to acquire manipulation skills from demonstrations, yet deploying a policy across tasks with heterogeneous dynamics remains challenging, as models tend to average over distinct behavioral modes present in the demonstrations. Mixture-of-Experts (MoE) architectures address this by activating specialized subnetworks, but requires meaningful skill decompositions for expert routing. We introduce Latent-Aligned Routing for Mixture of Experts (LAR-MoE), a two-stage framework that decouples unsupervised skill discovery from policy learning. In pre-training, we learn a joint latent representation between observations and future actions through student-teacher co-training. In a post-training stage, the expert routing is regularized to follow the structure of the learned latent space, preventing expert collapse while maintaining parameter efficiency. We evaluate LAR-MoE in simulation and on hardware. On the LIBERO benchmark, our method achieves a 95.2% average success rate with 150M parameters. On a surgical bowel grasping and retraction task, LAR-MoE matches a supervised MoE baseline without requiring any phase annotations, and transfers zero-shot to ex vivo porcine tissue. Our findings suggest that latent-aligned routing provides a principled alternative to supervised skill decomposition, enabling structured expert specialization from unlabeled demonstrations.

机器人模仿学习混合专家无监督技能分解

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。