用专家混合模型实现机器人装配技能端到端生成
ATG-MoE: Autoregressive trajectory generation with mixture-of-experts for assembly skill learning
- 融合视觉、语言与本体感知,直接生成动作轨迹
- 模拟中平均抓取成功率96.3%,整体成功率达91.8%
- 单模型支持多技能整合,适合工业装配场景
柔性制造需要能适应不断变化任务、物体和环境的机器人系统。传统机器人编程耗时且僵化,现有基于学习的装配方法常存在定位泛化能力弱、多阶段设计复杂、多技能集成有限等问题。本文提出ATG-MoE,一种面向装配技能学习的端到端自回归轨迹生成方法,结合专家混合架构。该方法建立从多模态输入(RGB-D观测、自然语言指令、机器人本体感知)到操作轨迹的闭环映射,融合多模态特征以理解场景与任务,采用自回归序列建模生成时间连贯轨迹,并通过专家混合结构实现统一的多技能学习。相比分离感知与控制或独立训练各技能的传统方法,ATG-MoE直接将视觉信息融入轨迹生成,支持高效多技能集成。在压力减压阀装配任务的8项代表性装配技能上进行训练与评估,实验表明其在仿真中表现优异,平均抓取成功率为96.3%,平均整体成功率为91.8%,并展现出强泛化能力和有效的多技能整合。真实世界实验进一步验证了其在多技能工业装配中的实用性。
原文摘要 · Abstract (English)
Flexible manufacturing requires robot systems that can adapt to constantly changing tasks, objects, and environments. However, traditional robot programming is labor-intensive and inflexible, while existing learning-based assembly methods often suffer from weak positional generalization, complex multi-stage designs, and limited multi-skill integration capability. To address these issues, this paper proposes ATG-MoE, an end-to-end autoregressive trajectory generation method with mixture of experts for assembly skill learning from demonstration. The proposed method establishes a closed-loop mapping from multi-modal inputs, including RGB-D observations, natural language instructions, and robot proprioception to manipulation trajectories. It integrates multi-modal feature fusion for scene and task understanding, autoregressive sequence modeling for temporally coherent trajectory generation, and a mixture-of-experts architecture for unified multi-skill learning. In contrast to conventional methods that separate visual perception and control or train different skills independently, ATG-MoE directly incorporates visual information into trajectory generation and supports efficient multi-skill integration within a single model. We train and evaluate the proposed method on eight representative assembly skills from a pressure-reducing valve assembly task. Experimental results show that ATG-MoE achieves strong overall performance in simulation, with an average grasp success rate of 96.3% and an average overall success rate of 91.8%, while also demonstrating strong generalization and effective multi-skill integration. Real-world experiments further verify its practicality for multi-skill industrial assembly. The project page can be found at https://hwh23.github.io/ATG-MoE
Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。