arXiv:2603.15265cs.RO2026-03被引 3

用稀疏专家模块提升机器人多任务操作的泛化能力

MoE-ACT: Scaling Multi-Task Bimanual Manipulation with Sparse Language-Conditioned Mixture-of-Experts Transformers

  • 在Transformer中引入稀疏专家模块,按任务动态激活不同专家
  • 仿真与真实双臂实验中成功率平均提升33%
  • 适合需要多任务语言控制的机器人操作研究者

让机器人在统一策略下完成多种任务是实现场景化智能的关键。然而,任务间分布外差异常导致严重干扰和负迁移。为此,我们提出轻量级多任务模仿学习框架MoE-ACT,将稀疏混合专家(MoE)模块融入ACT的Transformer编码器。MoE层将统一策略分解为可独立调用的专家组件,通过自适应激活实现潜在空间中多任务动作分布的自然解耦。解码时,特征线性调制(FiLM)动态调节动作标记,增强动作生成与任务指令的一致性。同时,多尺度交叉注意力使策略能同时关注低层与高层语义特征,提供丰富的视觉信息用于机器人操作。进一步融合文本信息,使系统从纯视觉模型转变为以视觉为中心、语言条件化的动作生成系统。在仿真与真实双臂设置中的实验表明,MoE-ACT相比基线模型平均成功率达33%提升,展现出更强的鲁棒性与泛化能力。

原文摘要 · Abstract (English)

The ability of robots to handle multiple tasks under a unified policy is critical for deploying embodied intelligence in real-world household and industrial applications. However, out-of-distribution variation across tasks often causes severe task interference and negative transfer when training general robotic policies. To address this challenge, we propose a lightweight multi-task imitation learning framework for bimanual manipulation, termed Mixture-of-Experts-Enhanced Action Chunking Transformer (MoE-ACT), which integrates sparse Mixture-of-Experts (MoE) modules into the Transformer encoder of ACT. The MoE layer decomposes a unified task policy into independently invoked expert components. Through adaptive activation, it naturally decouples multi-task action distributions in latent space. During decoding, Feature-wise Linear Modulation (FiLM) dynamically modulates action tokens to improve consistency between action generation and task instructions. In parallel, multi-scale cross-attention enables the policy to simultaneously focus on both low-level and high-level semantic features, providing rich visual information for robotic manipulation. We further incorporate textual information, transitioning the framework from a purely vision-based model to a vision-centric, language-conditioned action generation system. Experimental validation in both simulation and a real-world dual-arm setup shows that MoE-ACT substantially improves multi-task performance. Specifically, MoE-ACT outperforms vanilla ACT by an average of 33% in success rate. These results indicate that MoE-ACT provides stronger robustness and generalization in complex multi-task bimanual manipulation environments. Our open-source project page can be found at https://j3k7.github.io/MoE-ACT/.

多任务学习机器人操作语言控制专家混合

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。