用语义结构路由专家,让机器人更高效地完成组合式操作任务。
Semantically Structured Mixture-of-Experts for Compositional Robotic Manipulation

- 基于语义结构动态分配动作块给专用专家
- 在多任务基准上参数效率提升,新任务迁移效果好
- 适合需要高效泛化与可解释性的机器人控制场景
基于扩散模型的策略已成高精度机器人操作的新标准,但面临可扩展性瓶颈:高性能模型计算开销大,轻量替代品常难以跨多样化多任务环境泛化。混合专家(MoE)架构通过仅激活部分参数提供了效率提升路径,但现有路由机制依赖低层噪声或潜在统计特征,忽视了操作任务的组合特性,导致可复用行为被碎片化,降低可解释性与迁移能力。本文提出语义结构化混合专家扩散策略(SMoDP),将专家专长与语义任务结构对齐。SMoDP利用轻量级、推理时的技能预测器,由视觉-语言模型(VLMs)离线标注监督,将动作块路由至针对特定行为阶段优化的专家。为确保可靠分配,提出双对比对齐策略:跨模态上将多模态观测与语言定义的技能语义对齐(跨模态),并在视觉差异大但功能相似的行为间保持路由一致性(同模态)。该方法在多任务基准上显著优于代表性扩散与MoE基线,在参数效率上表现更优,并通过参数高效的微调实现对新任务的有效组合迁移。
原文摘要 · Abstract (English)
Diffusion-based policies have established a new standard for precise robotic manipulation but face a critical scalability bottleneck: high-performance models are computationally expensive, while lightweight alternatives often fail to generalize across diverse multi-task environments. Mixture-of-Experts (MoE) architectures offer a promising path to efficiency by activating only a subset of parameters. However, existing MoE routing mechanisms typically rely on low-level noise or latent statistics, ignoring the compositional nature of manipulation tasks. This can fragment reusable behaviors across experts, limiting interpretability and transferability. We introduce Semantically Structured Mixture-of-Experts Diffusion Policy (SMoDP) for compositional robotic manipulation, a framework that grounds expert specialization in semantic task structure. SMoDP leverages a lightweight, inference-time skill predictor, supervised by offline annotations from Vision-Language Models (VLMs), to route action chunks to experts specialized for specific behavioral phases. To ensure robust assignment, we propose a dual contrastive alignment strategy that grounds multi-modal observations in language-defined skill semantics (Inter-modal) while enforcing routing consistency across visually distinct but functionally related behaviors (Intra-modal). Our approach outperforms representative diffusion and MoE-based baselines on multi-task benchmarks with significantly improved parameter efficiency and demonstrates effective compositional transfer to novel tasks through parameter-efficient fine-tuning. Project website: https://deng-cy20.github.io/SMoDP/
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。