用专家混合模型让机器人自动拆解任务,生成可复用的底层动作单元。
Emergent Compositional Skills in Mixture-of-Experts VLAs

- 采用简化版专家混合(MoE)动作头,端到端学习组合策略
- 多个任务中重复使用同一专家,对应不同低层行为
- 无需预设分解规则,自动涌现可解释的模块化动作
我们研究如何从专家示范中端到端学习组合式机器人策略,无需预先设定任务分解或层级结构。探讨在采用简化混合专家(MoE)动作头的视觉语言代理(VLA)中,是否能自发地将任务分解为可复用、可解释的底层动作单元。结果表明,所学专家在多个任务间高度复用,并始终对应于质性上不同的低层行为,说明路由器隐式学会了高层序列规划,而专家则充当组合性基本单元。该MoE模型在任务性能上达到单体基线水平,同时展现出有意义的专家专业化,朝着仅由数据驱动的模块化、可解释机器人策略迈出一步。
原文摘要 · Abstract (English)
We consider the problem of learning compositional robot policies end-to-end from expert demonstrations, without any pre-specified notion of task decomposition or hierarchy. We ask whether a VLA trained with a simplified Mixture-of-Experts (MoE) action head can emergently learn to decompose tasks into reusable, interpretable primitives. We find that learned experts are heavily reused across tasks and consistently correspond to qualitatively distinct low-level behaviors, suggesting that the router implicitly learns to perform high-level sequencing while experts serve as compositional primitives. Our MoE matches the task performance of a monolithic baseline while demonstrating meaningful expert specialization, a step toward modular, interpretable robot policies that emerge from data alone.
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。