arXiv:2410.02475cs.ROcs.LG2024-10ICLR被引 18

用专家混合模型实现高效通用抓取,12小时搞定3200种物体。

Efficient Residual Learning with Mixture-of-Experts for Universal Dexterous Grasping

  • 基于残差学习与专家混合架构,提升多任务抓取效率。
  • 在3200个物体上达88.8%成功率,无未见物体泛化差距。
  • 适合机器人抓取、强化学习多任务训练场景研究者。

跨多样化物体的通用灵巧抓取是机器人学习中的基础而严峻挑战。现有基于强化学习(RL)的方法在大规模物体数据集上训练策略时面临关键局限:多任务学习需复杂课程设计,且对未见物体泛化能力有限。为此,我们提出ResDex,一种将残差策略学习与混合专家(MoE)框架结合的新方法。ResDex采用几何无关的基策略,在单个物体上高效获取,具备广泛泛化至未见物体的能力。其MoE框架融合多个基策略,支持多样抓取风格适配不同物体。通过学习残差动作及组合基策略的权重,ResDex实现了高效的多任务强化学习。在包含3200个物体的DexGraspNet数据集上,ResDex达到88.8%的成功率,对未见物体无泛化性能下降,并仅用单张GPU在12小时内完成全部任务训练。

原文摘要 · Abstract (English)

Universal dexterous grasping across diverse objects presents a fundamental yet formidable challenge in robot learning. Existing approaches using reinforcement learning (RL) to develop policies on extensive object datasets face critical limitations, including complex curriculum design for multi-task learning and limited generalization to unseen objects. To overcome these challenges, we introduce ResDex, a novel approach that integrates residual policy learning with a mixture-of-experts (MoE) framework. ResDex is distinguished by its use of geometry-unaware base policies that are efficiently acquired on individual objects and capable of generalizing across a wide range of unseen objects. Our MoE framework incorporates several base policies to facilitate diverse grasping styles suitable for various objects. By learning residual actions alongside weights that combine these base policies, ResDex enables efficient multi-task RL for universal dexterous grasping. ResDex achieves state-of-the-art performance on the DexGraspNet dataset comprising 3,200 objects with an 88.8% success rate. It exhibits no generalization gap with unseen objects and demonstrates superior training efficiency, mastering all tasks within only 12 hours on a single GPU.

灵巧抓取强化学习专家混合机器人

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。