arXiv:2501.03606cs.ROcs.CV2025-01被引 6

通过多模态预训练实现双手灵巧操作,提升复杂任务成功率。

VTAO-BiManip: Masked Visual-Tactile-Action Pre-training with Object Understanding for Bimanual Dexterous Manipulation

  • 融合视觉、触觉与动作数据进行掩码建模,增强双臂协调能力。
  • 在拧瓶盖任务中成功率超现有方法20%以上,模拟与真实环境均有效。
  • 适合需要多技能学习的机器人双手操作研究者参考。

双手灵巧操作因每只手自由度高且需协同而面临挑战。现有单手操作方法依赖人类示范引导强化学习,但难以泛化至涉及多种子技能的复杂双臂任务。本文提出VTAO-BiManip框架,结合视觉-触觉-动作预训练与物体理解,支持课程化强化学习,实现类人双手操作。通过引入手部运动数据,相比仅用二值触觉反馈,能更有效地指导双臂协调。预训练模型利用掩码多模态输入预测未来动作、物体位姿与尺寸,促进跨模态正则化。针对多技能学习难题,设计两阶段课程强化学习策略以稳定训练。在瓶盖拧开任务上评估,该方法在仿真与真实环境中均表现优异,成功率超过现有视觉-触觉预训练方法20%以上。

原文摘要 · Abstract (English)

Bimanual dexterous manipulation remains significant challenges in robotics due to the high DoFs of each hand and their coordination. Existing single-hand manipulation techniques often leverage human demonstrations to guide RL methods but fail to generalize to complex bimanual tasks involving multiple sub-skills. In this paper, we introduce VTAO-BiManip, a novel framework that combines visual-tactile-action pretraining with object understanding to facilitate curriculum RL to enable human-like bimanual manipulation. We improve prior learning by incorporating hand motion data, providing more effective guidance for dual-hand coordination than binary tactile feedback. Our pretraining model predicts future actions as well as object pose and size using masked multimodal inputs, facilitating cross-modal regularization. To address the multi-skill learning challenge, we introduce a two-stage curriculum RL approach to stabilize training. We evaluate our method on a bottle-cap unscrewing task, demonstrating its effectiveness in both simulated and real-world environments. Our approach achieves a success rate that surpasses existing visual-tactile pretraining methods by over 20%.

双手操作多模态预训练强化学习机器人灵巧操作

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。