arXiv:2412.06779cs.ROcs.AI2024-12ICCV被引 28

用少量实操数据,让单手策略学会复杂双手操作。

AnyBimanual: Transferring Unimanual Policy for General Bimanual Manipulation

  • 用技能调度器融合预训练单手策略,生成双手动作指令。
  • 仿真任务成功率提升12.67%,真实场景平均成功率84.62%。
  • 适合想快速部署通用双手操作的机器人研究者和工程师。

通用语言控制的双手操作在家庭服务与工业装配中至关重要,但因动作空间高维,收集双手操作数据成本高昂,传统方法难以应对。相比之下,近期单手策略凭借大规模参数与数据训练展现出优异泛化能力,可为双手系统提供可复用的操作知识。为此,我们提出一种即插即用的方法 AnyBimanual,仅需少量双手示范即可将预训练单手策略迁移到通用双手操作任务。具体而言,我们引入技能管理器,动态调度从预训练单手策略中提取的技能表征,通过任务导向补偿线性组合技能基元,以表示双手操作指令。为缓解单手与双手系统间的观测差异,设计视觉对齐模块,生成工作区软掩码,对齐视觉嵌入特征。在 RLBench2 的 12 个仿真任务上,成功率达 12.67% 的显著提升;9 项真实世界任务实验进一步验证其实用性,平均成功率达 84.62%。

原文摘要 · Abstract (English)

Performing general language-conditioned bimanual manipulation tasks is of great importance for many applications ranging from household service to industrial assembly. However, collecting bimanual manipulation data is expensive due to the high-dimensional action space, which poses challenges for conventional methods to handle general bimanual manipulation tasks. In contrast, unimanual policy has recently demonstrated impressive generalizability across a wide range of tasks because of scaled model parameters and training data, which can provide sharable manipulation knowledge for bimanual systems. To this end, we propose a plug-and-play method named AnyBimanual, which transfers pre-trained unimanual policy to general bimanual manipulation policy with few bimanual demonstrations. Specifically, we first introduce a skill manager to dynamically schedule the skill representations discovered from pre-trained unimanual policy for bimanual manipulation tasks, which linearly combines skill primitives with task-oriented compensation to represent the bimanual manipulation instruction. To mitigate the observation discrepancy between unimanual and bimanual systems, we present a visual aligner to generate soft masks for visual embedding of the workspace, which aims to align visual input of unimanual policy model for each arm with those during pretraining stage. AnyBimanual shows superiority on 12 simulated tasks from RLBench2 with a sizable 12.67% improvement in success rate over previous methods. Experiments on 9 real-world tasks further verify its practicality with an average success rate of 84.62%.

机器人策略迁移双手操作强化学习

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。