MoSE让机器人像人一样分步学技能,效率更高。
MoSE: Skill-by-Skill Mixture-of-Experts Learning for Embodied Autonomous Machines
- 按技能分步调度专家,模仿人类学习方式。
- 参数少于40%却在驾驶和机器人任务上超越大模型。
- 单次前向计算融合感知、规划等多任务,零额外开销。
为应对智能体自主系统对更高效推理与学习的需求,本文提出一种新型分层专家混合(MoE)方法——MoSE,显著提升具身智能体的推理与学习效率。通用MoE需大量训练数据与复杂优化,难以应用于自动驾驶(AD)和机器人操作等场景。为此,我们设计了一种以技能为导向的MoE架构,模拟人类分步学习过程。通过定义并标注具体技能,引入技能导向的路由机制,使专家能识别不同场景下的必要能力,实现逐技能学习。为匹配人类多步推理及端到端驾驶模型的规划需求,构建了层次化技能数据集,并预训练路由模块以促进分步思考。不同于多轮对话,MoSE在单次前向传播中集成感知-预测-规划(如自动驾驶)或高层-低层规划(如机器人)等辅助任务,不增加任何计算成本。仅使用不到30亿稀疏激活参数,模型便具备更丰富的专业知识,在自动驾驶边缘案例推理和机器人推理任务上表现优于参数量超过40%的模型。
原文摘要 · Abstract (English)
To meet the growing demand for smarter, faster, and more efficient embodied AI solutions, we introduce a novel Mixture-of-Expert (MoE) method that significantly boosts reasoning and learning efficiency for embodied autonomous systems. General MoE models demand extensive training data and complex optimization, which limits their applicability in embodied AI such as autonomous driving (AD) and robotic manipulation. In this work, we propose a skill-oriented MoE called MoSE, which mimics the human learning and reasoning process skill-by-skill, step-by-step. We introduce a skill-oriented routing mechanism that begins with defining and annotating specific skills, enabling experts to identify the necessary competencies for various scenarios and reasoning tasks, thereby facilitating skill-by-skill learning. To better align with multi-step planning in human reasoning and in end-to-end driving models, we build a hierarchical skill dataset and pretrain the router to encourage the model to think step-by-step. Unlike other multi-round dialogues, MoSE integrates valuable auxiliary tasks (e.g. perception-prediction-planning for AD, and high-level and low-level planning for robots) in one single forward process without introducing any extra computational cost. With less than 3B sparsely activated parameters, our model effectively grows more diverse expertise and outperforms models on both AD corner-case reasoning tasks and robot reasoning tasks with less than 40% of the parameters.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。