arXiv:2410.07403cs.RO2024-10被引 3

混合方法提升复杂工具操作的长期任务执行能力

On the Feasibility of A Mixed-Method Approach for Solving Long Horizon Task-Oriented Dexterous Manipulation

  • 结合模仿学习、强化学习与模型控制,分阶段解决工具操作难题
  • 多子任务表现优于单一强化学习,仿真中成功率显著提升
  • 适合需要精细操控的机器人抓取与真实场景迁移研究者

使用灵巧手在现实世界中进行工具的手中操作是一个研究较少的问题。相比常见的立方体或圆柱体,工具具有更复杂的几何形状和更大尺寸,且任务导向的手中工具操作需按顺序完成多个子任务:触达工具、抓取、在手中重新定向(可能需重新抓握)以达到合适的使用姿态,并将工具运送到目标位置。现有研究多采用强化学习分别学习各子任务并组合策略完成长时程任务,但单一方法难以适应所有子任务,尤其对多指灵巧手操控复杂物体时更为明显。本文提出一种混合方法,融合模仿学习、强化学习与基于模型的控制。同时引入基于强化学习的教师-学生框架,将真实数据融入离线训练。实验表明,该方法在不同子任务及长时程任务中均优于传统强化学习,在仿真中表现更优;最终成功实现向真实世界的迁移。

原文摘要 · Abstract (English)

In-hand manipulation of tools using dexterous hands in real-world is an underexplored problem in the literature. In addition to more complex geometry and larger size of the tools compared to more commonly used objects like cubes or cylinders, task oriented in-hand tool manipulation involves many sub-tasks to be performed sequentially. This may involve reaching to the tool, picking it up, reorienting it in hand with or without regrasping to reach to a desired final grasp appropriate for the tool usage, and carrying the tool to the desired pose. Research on long-horizon manipulation using dexterous hands is rather limited and the existing work focus on learning the individual sub-tasks using a method like reinforcement learning (RL) and combine the policies for different subtasks to perform a long horizon task. However, in general a single method may not be the best for all the sub-tasks, and this can be more pronounced when dealing with multi-fingered hands manipulating objects with complex geometry like tools. In this paper, we investigate the use of a mixed-method approach to solve for the long-horizon task of tool usage and we use imitation learning, reinforcement learning and model based control. We also discuss a new RL-based teacher-student framework that combines real world data into offline training. We show that our proposed approach for each subtask outperforms the commonly adopted reinforcement learning approach across different subtasks and in performing the long horizon task in simulation. Finally we show the successful transferability to real world.

灵巧操作强化学习仿真到真实

Thank you to arXiv for use of its open access interoperability. PaperDance 不是 arXiv 官方产品;中文卡片由大模型生成,请以原文为准。